03API change analysis
Callsite Impact
In short
- What it does
- Predicts which parts of an app will break when an outside service it relies on releases a new version of its API.
- Why it matters
- A tool that compares API documents can list hundreds of changes, but not which of them break your own code.
- What I built
- Built a pipeline that writes calls against the old API, lets the compiler mark which break on the new one, and predicts them without seeing its answer.
- Key result
- On the APIs it was developed on it scored 0.963 (F1); on APIs held back until the end, 0.328 — no false alarms, but only about one breakage in five caught. The second number is the honest one.
Overview
Predicts which call sites break under an OpenAPI change, with compiler-verified labels as ground truth and a held-out slice frozen before it was scored.
- Problem
- A spec differ can report hundreds of breaking changes between two API versions but cannot say which of your call sites break — it compares documents and knows nothing about your code. Adyen BalancePlatform v1 → v2: 296 reported changes; the compiler breaks 33 of 133 call sites.
- Built
- A pipeline that generates call sites from revision A, keeps only those that compile, lets the TypeScript compiler label what breaks under revision B, and predicts those breakages from the oasdiff change set and a parsed view of the source — never from compiler output.
- Hard part
- Keeping the author out of the answer key: the generator never sees revision B or the diff, a failing call site is discarded rather than repaired, and a guard fails the build if the classifier imports the oracle.
Evidence
F1 score on the development corpus vs a slice held out and frozen before scoring; on unseen APIs it misses four breakages in five
The development F1 is a post-selection number; the held-out one is the unbiased estimate.
held-out precision: zero false positives across 1,077 clean call sites
At a held-out recall of 0.196.
compiler-verified breakages across 3 vendors in the release corpus
call sites a naive 'changed operation' baseline flags to catch the same 94
Development vs held-out
| Measure | Development | Held-out slice |
|---|---|---|
| Spec pairs / vendors | 18 / 3 | 15 / 3 |
| Admitted call sites | 1,338 | 1,342 |
| Compiler-verified breakages | 94 | 265 |
| Precision | 0.958 | 1.000 |
| Recall | 0.968 | 0.196 |
| F1 | 0.963 | 0.328 |
Frozen in git before it was measured, scored once, and nothing changed afterwards. On unseen services the system is never wrong when it speaks, and misses four breakages in five.
Source: README · held-out slice, frozen at 9378d58
Architecture
Ground truth
- A01Input
Vendor spec A
- A02Code
Generate call sites
from revision A and a seed
- A03Gate
tsc against A
admit only clean call sites
- A04Store
tsc against B
compiler labels = ground truth
System under test
- B01Input
Specs A + B
- B02Code
oasdiff change set
- B03Code
Parsed call-site source
no type checker
- B04Gate
Rule table
IMPACTED · UNAFFECTED · UNKNOWN
Neither the answer key nor the call sites it grades are written by a person.
Verdicts are graded against the compiler's labels. The classification package is forbidden from importing the oracle.
- InputArrives from outside the system
- CodeDeterministic code
- GateDecides whether work proceeds
- StoreDurable state
Engineering notes
Why it fails on unseen services, traced to one character
39% of the held-out misses are on operations the differ reported no change for at all. oasdiff normalises path-parameter names — {EmployeeId} and {EmployeeID} are the same endpoint to it — while openapi-typescript keys paths on the literal string, so every call site on that path breaks. The tool's recall is capped by the differ's recall, and the gap is silent.
It is not fixed. Fixing it against the slice that revealed it would turn the only unbiased number in the repository into a second development number.
Three verdicts, and why the third is not a cop-out
UNKNOWN covers changes the type system cannot express — a decreased maxLength is a real breaking change, but no call site can be made to fail on it, so a clean compile is not evidence of safety. The abstention rate is published: 2.2% of (call site × change) pairs on development, 1.8% on the held-out slice.
A kill criterion sensitive to budget, and published that way
The predeclared criterion was at least 60 compiler-verified breakages across at least 3 vendors. At the harness default budget the same corpus yields 52 and fails; the canonical release budget yields 94 and passes. Both are published side by side, and F1 moves by 0.002 while the admitted corpus grows 2.3×, from 854 to 1,931 call sites.
Screens


Limitations
As the project states them. Read these before relying on any number above.
- Held-out recall is 0.196: on unseen services the tool misses four breakages in five.
- The largest group of held-out misses (84 of 213) is on operations the differ reported no change for; the cause found, path-parameter renames it normalises away, is deliberately not fixed against the spent slice.
- The kill-criterion count depends on the generation budget; the predictive metrics do not.
- No model is used anywhere, and no retrieval layer was built.
Facts and stack
- Vendors
- Adyen, Twilio, Xero — MIT-licensed specs
- Release corpus
- 18 revision pairs, 1,338 call sites
- Abstention
- 2.2% of (call site × change) pairs; 1.8% held out
Stack
- Python 3.12
- TypeScript compiler
- openapi-typescript
- oasdiff
- FastAPI (local, read-only)
- Next.js 15
- Vercel
Skills shown
- API change-impact analysis
- OpenAPI
- Evaluation design
- Held-out testing
- Static analysis