Skip to content
Tahir Aslanli

05Entity resolution

Counterparty Resolver

Complete · frozenDeployed / live

Free demo — may take ~1 min to wake.

In short

What it does
Decides whether two company records describe the same business, and leaves every merge for a person to approve.
Why it matters
Wrongly merging two companies corrupts where payments go, and missing a duplicate leaves the same company on file twice.
What I built
Built clear, explainable matching rules over public company-registry data, a person approving every merge, and a merge history that can be undone exactly.
Key result
Precision of 0.9980 on 10,532 development pairs — six false merges against seven for the best baseline — though a simple nine-line rule still scores higher overall (F1 0.7607 against 0.7303).

Overview

Entity resolution with deterministic rules over interpretable features, a human approving every merge, and a merge ledger that reverses to the byte — evaluated on real GLEIF adjudications.

Problem
A wrong merge corrupts payment routing for days; a missed duplicate costs a duplicate. So the positive labels here are not the author's: they are duplicate adjudications recorded by LEI Issuing Organisations before this project existed.
Built
A resolution layer over fixed legacy schemas that may not change: deterministic blocking, ordered rules that decide MATCH, REVIEW or NO_MATCH, human approval of every merge, and an append-only merge ledger that reverses to the byte — evaluated on registrar duplicate adjudications from the public GLEIF dataset, the only data the corpus and the live demo hold.
Hard part
Brownfield constraints: the source schemas may not change, so everything is additive — including an expand, backfill and contract migration that refuses to drop a column while any row would lose its approver.

Evidence

0.9980

precision over 10,532 development pairs — six false merges, against seven for the best baseline; on F1 the nine-line identifier_first baseline still wins, 0.7607 to 0.7303

Held out: 0.9986 with one false merge, or 0.9972 with development-only priors. The project quotes the development figure: the hold-out's negatives are easier, so its absolute precision is not comparable.

< 0.005

movement of the system's precision, recall and F1 between development and the hold-out

12,984

labelled pairs: GLEIF duplicate adjudications plus mined hard negatives

0.9138

candidate recall at a 0.9955 reduction ratio

Measured over a record pool made up of known duplicates; a real master file would give a different ratio.

Held-out precision and F1, against the three predeclared baselines

PrecisionF1
Held-out precision and F1, against the three predeclared baselines
MethodPrecisionF1
exact_normalized_name0.99670.6572
fuzzy_name_only_0.900.88900.7488
identifier_first0.98170.7538
system0.99860.7344

The kill test was predeclared on precision, and the system wins there. On F1, identifier_first and the fuzzy baseline both score higher; the project publishes both, and notes that the fuzzy row is not a fair measurement.

Source: artifacts/evaluation.json · hold-out scored once at b67b83e

Architecture

  1. 01Input

    GLEIF records

    fixed legacy schema, read through a view

  2. 02Code

    Blocking

    candidate reduction

  3. 03Gate

    Registrar number

    same authority and number, on at most two records → MATCH

  4. 04Gate

    Series / vintage designator

    differing designator → two products

  5. 05Code

    Weighted features

    decides REVIEW or NO_MATCH — never MATCH

  6. 06Human

    Human approval

  7. 07Store

    Append-only merge ledger

    reversible, including source links

What a registrar wrote is read first; what two strings look like is read afterwards.

No model is in the decision path. Exact agreement of the whole canonicalised name also yields MATCH.

  • InputArrives from outside the system
  • CodeDeterministic code
  • HumanA person decides
  • GateDecides whether work proceeds
  • StoreDurable state

Engineering notes

Positive labels nobody here wrote

The labels are duplicate adjudications recorded by LEI Issuing Organisations in the public GLEIF dataset (CC0), paired with an equal number of similarity-mined hard negatives. The split is by entity cluster, not by pair, so no entity appears on both sides. The entity split was committed before any rule existed; the hold-out's negatives were later redrawn within the held-out side, and the hold-out was scored once.

Why the weighted score never asserts a match

A MATCH band at 0.86 decided 234 development pairs, 83 of them wrongly — a rule wrong a third of the time is a queue, not an auto-merge. Removing the band cost 0.029 recall and removed 83 of 89 false merges. The whole sweep is published so the choice is auditable.

An identifier only counts when it identifies: every Allianz fund at one registration authority shared a single number, and the first rule merged a small-cap equity fund with a bond fund — 40 false merges — until agreement required the number to appear on at most two records.

A merge ledger that can be unpicked

Append-only in the database, enforced by triggers that refuse UPDATE and DELETE. Idempotent on a unique key, so a retry after a timeout is one merge. The unmerge test compares every table except the ledger itself — which must hold exactly the merge and its reversal — before the merge and after the reversal, because asserting that the resolved entity disappeared would pass while leaving source links pointing nowhere.

Screens

Pair evidence screen listing every feature, what it compared and what it contributedPair evidence screen listing every feature, what it compared and what it contributed
One pair: every feature, what it compared, and what it contributed.Captured from the live deployment
Review queue of pairs the resolver declined to decide, highest score firstReview queue of pairs the resolver declined to decide, highest score first
The review queue: what the resolver declined to decide, highest score first.Captured from the live deployment

Limitations

As the project states them. Read these before relying on any number above.

  • identifier_first beats the system on F1: 0.7607 against 0.7303 on development.
  • The precision margin over the best baseline, exact_normalized_name, is one false merge — six against seven on development.
  • The hold-out's negatives are measurably easier than development's, so its absolute precision is not comparable.
  • Prior names are carried but not consulted: 324 missed development duplicates have an exact alias match.
  • The REVIEW band is measured but not adjudicated by a model, so no cost comparison between arms is published.
  • The live demo is read-only: every write is refused with 403, so a merge cannot be approved there.

Facts and stack

Records
12,853
Review band
21.2% of development pairs handed to a person
Planted breaches
25 guards, each planted and caught

Stack

  • Python
  • FastAPI
  • SQLite
  • rapidfuzz
  • GLEIF (CC0)
  • Docker
  • Render

Skills shown

  • Entity resolution
  • Evaluation vs baselines
  • Brownfield migration
  • Append-only ledger
  • Human approval