05Entity resolution
Counterparty Resolver
Free demo — may take ~1 min to wake.
In short
- What it does
- Decides whether two company records describe the same business, and leaves every merge for a person to approve.
- Why it matters
- Wrongly merging two companies corrupts where payments go, and missing a duplicate leaves the same company on file twice.
- What I built
- Built clear, explainable matching rules over public company-registry data, a person approving every merge, and a merge history that can be undone exactly.
- Key result
- Precision of 0.9980 on 10,532 development pairs — six false merges against seven for the best baseline — though a simple nine-line rule still scores higher overall (F1 0.7607 against 0.7303).
Overview
Entity resolution with deterministic rules over interpretable features, a human approving every merge, and a merge ledger that reverses to the byte — evaluated on real GLEIF adjudications.
- Problem
- A wrong merge corrupts payment routing for days; a missed duplicate costs a duplicate. So the positive labels here are not the author's: they are duplicate adjudications recorded by LEI Issuing Organisations before this project existed.
- Built
- A resolution layer over fixed legacy schemas that may not change: deterministic blocking, ordered rules that decide MATCH, REVIEW or NO_MATCH, human approval of every merge, and an append-only merge ledger that reverses to the byte — evaluated on registrar duplicate adjudications from the public GLEIF dataset, the only data the corpus and the live demo hold.
- Hard part
- Brownfield constraints: the source schemas may not change, so everything is additive — including an expand, backfill and contract migration that refuses to drop a column while any row would lose its approver.
Evidence
precision over 10,532 development pairs — six false merges, against seven for the best baseline; on F1 the nine-line identifier_first baseline still wins, 0.7607 to 0.7303
Held out: 0.9986 with one false merge, or 0.9972 with development-only priors. The project quotes the development figure: the hold-out's negatives are easier, so its absolute precision is not comparable.
movement of the system's precision, recall and F1 between development and the hold-out
labelled pairs: GLEIF duplicate adjudications plus mined hard negatives
candidate recall at a 0.9955 reduction ratio
Measured over a record pool made up of known duplicates; a real master file would give a different ratio.
Held-out precision and F1, against the three predeclared baselines
| Method | Precision | F1 |
|---|---|---|
| exact_normalized_name | 0.9967 | 0.6572 |
| fuzzy_name_only_0.90 | 0.8890 | 0.7488 |
| identifier_first | 0.9817 | 0.7538 |
| system | 0.9986 | 0.7344 |
The kill test was predeclared on precision, and the system wins there. On F1, identifier_first and the fuzzy baseline both score higher; the project publishes both, and notes that the fuzzy row is not a fair measurement.
Source: artifacts/evaluation.json · hold-out scored once at b67b83e
Architecture
- 01Input
GLEIF records
fixed legacy schema, read through a view
- 02Code
Blocking
candidate reduction
- 03Gate
Registrar number
same authority and number, on at most two records → MATCH
- 04Gate
Series / vintage designator
differing designator → two products
- 05Code
Weighted features
decides REVIEW or NO_MATCH — never MATCH
- 06Human
Human approval
- 07Store
Append-only merge ledger
reversible, including source links
What a registrar wrote is read first; what two strings look like is read afterwards.
No model is in the decision path. Exact agreement of the whole canonicalised name also yields MATCH.
- InputArrives from outside the system
- CodeDeterministic code
- HumanA person decides
- GateDecides whether work proceeds
- StoreDurable state
Engineering notes
Positive labels nobody here wrote
The labels are duplicate adjudications recorded by LEI Issuing Organisations in the public GLEIF dataset (CC0), paired with an equal number of similarity-mined hard negatives. The split is by entity cluster, not by pair, so no entity appears on both sides. The entity split was committed before any rule existed; the hold-out's negatives were later redrawn within the held-out side, and the hold-out was scored once.
Why the weighted score never asserts a match
A MATCH band at 0.86 decided 234 development pairs, 83 of them wrongly — a rule wrong a third of the time is a queue, not an auto-merge. Removing the band cost 0.029 recall and removed 83 of 89 false merges. The whole sweep is published so the choice is auditable.
An identifier only counts when it identifies: every Allianz fund at one registration authority shared a single number, and the first rule merged a small-cap equity fund with a bond fund — 40 false merges — until agreement required the number to appear on at most two records.
A merge ledger that can be unpicked
Append-only in the database, enforced by triggers that refuse UPDATE and DELETE. Idempotent on a unique key, so a retry after a timeout is one merge. The unmerge test compares every table except the ledger itself — which must hold exactly the merge and its reversal — before the merge and after the reversal, because asserting that the resolved entity disappeared would pass while leaving source links pointing nowhere.
Screens




Limitations
As the project states them. Read these before relying on any number above.
- identifier_first beats the system on F1: 0.7607 against 0.7303 on development.
- The precision margin over the best baseline, exact_normalized_name, is one false merge — six against seven on development.
- The hold-out's negatives are measurably easier than development's, so its absolute precision is not comparable.
- Prior names are carried but not consulted: 324 missed development duplicates have an exact alias match.
- The REVIEW band is measured but not adjudicated by a model, so no cost comparison between arms is published.
- The live demo is read-only: every write is refused with 403, so a merge cannot be approved there.
Facts and stack
- Records
- 12,853
- Review band
- 21.2% of development pairs handed to a person
- Planted breaches
- 25 guards, each planted and caught
Stack
- Python
- FastAPI
- SQLite
- rapidfuzz
- GLEIF (CC0)
- Docker
- Render
Skills shown
- Entity resolution
- Evaluation vs baselines
- Brownfield migration
- Append-only ledger
- Human approval