Skip to content
Tahir Aslanli

03API change analysis

Callsite Impact

Complete · frozenDeployed / live

In short

What it does
Predicts which parts of an app will break when an outside service it relies on releases a new version of its API.
Why it matters
A tool that compares API documents can list hundreds of changes, but not which of them break your own code.
What I built
Built a pipeline that writes calls against the old API, lets the compiler mark which break on the new one, and predicts them without seeing its answer.
Key result
On the APIs it was developed on it scored 0.963 (F1); on APIs held back until the end, 0.328 — no false alarms, but only about one breakage in five caught. The second number is the honest one.

Overview

Predicts which call sites break under an OpenAPI change, with compiler-verified labels as ground truth and a held-out slice frozen before it was scored.

Problem
A spec differ can report hundreds of breaking changes between two API versions but cannot say which of your call sites break — it compares documents and knows nothing about your code. Adyen BalancePlatform v1 → v2: 296 reported changes; the compiler breaks 33 of 133 call sites.
Built
A pipeline that generates call sites from revision A, keeps only those that compile, lets the TypeScript compiler label what breaks under revision B, and predicts those breakages from the oasdiff change set and a parsed view of the source — never from compiler output.
Hard part
Keeping the author out of the answer key: the generator never sees revision B or the diff, a failing call site is discarded rather than repaired, and a guard fails the build if the classifier imports the oracle.

Evidence

0.963 → 0.328Negative result

F1 score on the development corpus vs a slice held out and frozen before scoring; on unseen APIs it misses four breakages in five

The development F1 is a post-selection number; the held-out one is the unbiased estimate.

1.000

held-out precision: zero false positives across 1,077 clean call sites

At a held-out recall of 0.196.

94

compiler-verified breakages across 3 vendors in the release corpus

895

call sites a naive 'changed operation' baseline flags to catch the same 94

Development vs held-out

Corpus size and scores on the development corpus and on the held-out slice.
MeasureDevelopmentHeld-out slice
Spec pairs / vendors18 / 315 / 3
Admitted call sites1,3381,342
Compiler-verified breakages94265
Precision0.9581.000
Recall0.9680.196
F10.9630.328

Frozen in git before it was measured, scored once, and nothing changed afterwards. On unseen services the system is never wrong when it speaks, and misses four breakages in five.

Source: README · held-out slice, frozen at 9378d58

Architecture

Ground truth

  1. A01Input

    Vendor spec A

  2. A02Code

    Generate call sites

    from revision A and a seed

  3. A03Gate

    tsc against A

    admit only clean call sites

  4. A04Store

    tsc against B

    compiler labels = ground truth

System under test

  1. B01Input

    Specs A + B

  2. B02Code

    oasdiff change set

  3. B03Code

    Parsed call-site source

    no type checker

  4. B04Gate

    Rule table

    IMPACTED · UNAFFECTED · UNKNOWN

Neither the answer key nor the call sites it grades are written by a person.

Verdicts are graded against the compiler's labels. The classification package is forbidden from importing the oracle.

  • InputArrives from outside the system
  • CodeDeterministic code
  • GateDecides whether work proceeds
  • StoreDurable state

Engineering notes

Why it fails on unseen services, traced to one character

39% of the held-out misses are on operations the differ reported no change for at all. oasdiff normalises path-parameter names — {EmployeeId} and {EmployeeID} are the same endpoint to it — while openapi-typescript keys paths on the literal string, so every call site on that path breaks. The tool's recall is capped by the differ's recall, and the gap is silent.

It is not fixed. Fixing it against the slice that revealed it would turn the only unbiased number in the repository into a second development number.

Three verdicts, and why the third is not a cop-out

UNKNOWN covers changes the type system cannot express — a decreased maxLength is a real breaking change, but no call site can be made to fail on it, so a clean compile is not evidence of safety. The abstention rate is published: 2.2% of (call site × change) pairs on development, 1.8% on the held-out slice.

A kill criterion sensitive to budget, and published that way

The predeclared criterion was at least 60 compiler-verified breakages across at least 3 vendors. At the harness default budget the same corpus yields 52 and fails; the canonical release budget yields 94 and passes. Both are published side by side, and F1 moves by 0.002 while the admitted corpus grows 2.3×, from 854 to 1,931 call sites.

Screens

The measured result page comparing development and held-out scores
The measured result, including the held-out slice — captured from the deployed site.Captured from the live deployment
Table of call sites with the compiler, system and baseline verdicts side by side
The compiler, the system and the baseline on the same call sites.Captured from the live deployment

Limitations

As the project states them. Read these before relying on any number above.

  • Held-out recall is 0.196: on unseen services the tool misses four breakages in five.
  • The largest group of held-out misses (84 of 213) is on operations the differ reported no change for; the cause found, path-parameter renames it normalises away, is deliberately not fixed against the spent slice.
  • The kill-criterion count depends on the generation budget; the predictive metrics do not.
  • No model is used anywhere, and no retrieval layer was built.

Facts and stack

Vendors
Adyen, Twilio, Xero — MIT-licensed specs
Release corpus
18 revision pairs, 1,338 call sites
Abstention
2.2% of (call site × change) pairs; 1.8% held out

Stack

  • Python 3.12
  • TypeScript compiler
  • openapi-typescript
  • oasdiff
  • FastAPI (local, read-only)
  • Next.js 15
  • Vercel

Skills shown

  • API change-impact analysis
  • OpenAPI
  • Evaluation design
  • Held-out testing
  • Static analysis