01Finance operations
Ledger Exception Control Plane
In short
- What it does
- Resolves mismatches between payment records and a company's books, and is designed so a retried request never posts the same correction twice.
- Why it matters
- A duplicated or wrong correction quietly distorts reported revenue, and is usually found months later.
- What I built
- Built automatic matching, AI that suggests fixes but never amounts, and a person's approval before anything is posted.
- Key result
- In 21 of 21 failure-test runs no correction was posted twice, while a simpler version double-posted in 5 of 7 scenarios. In one run against a live model, all 177 of its wrong suggestions were refused; the public demo uses a labelled stand-in.
Overview
A finance-operations control plane: deterministic matching clears the bulk, a model proposes a treatment code for the residual, a human authorises it, and code computes the amount.
- Problem
- Payment-provider settlement files never fully match the general ledger, and the residual is resolved by hand at month-end. The expensive failures are silent: an adjustment posted twice because a retry fired, or posted to the wrong account, surfaces months later in reported revenue.
- Built
- Ingestion with batch quarantine, deterministic matching with per-currency tolerance bands, one classified exception per residual with an evidence pack, a model confined to a closed enum of treatment codes, role-separated human approval, a pure Decimal amount calculator, and a transactional outbox with bounded retry, a dead-letter queue and a recovery path for ambiguous outcomes.
- Hard part
- A send whose outcome nobody knows: the ledger may have committed the posting and the response was lost. The system records it as UNKNOWN, never retries it on the assumption it failed, and follows the adapter's declared capability: a bounded re-send under the same operation id where the ledger enforces the key, reconciliation by query where it can be queried, otherwise an operator.
Evidence
failure-scenario runs (7 scenarios × 3 ledger configurations) in which no adjustment was applied twice, counted by the simulated ledger itself
Where the ledger enforces the key, the simulated ledger does the suppressing: those runs show the dispatcher behaving correctly given an enforcing ledger, not that any real ledger enforces anything.
failure scenarios in which a deliberately naive version posted the same adjustment twice
The baseline is a legitimate implementation missing specific safeguards, documented omission by omission.
results, across both versions, that matched expectations declared before the run
live model accuracy over 247 answered records; answering “escalate” every time would score 85.6%
The model often proposed a treatment where escalation was correct. The account policy refuses all 177 such answers, and human approval stands in front of the one that would have priced. On the 36 priceable records it scored 97.2%.
Chaos suite · applied postings
Each number is how many times the ledger applied one adjustment: 1 is correct, 2 is a double posting.
Crash before commit
main111naive2 (posted twice)2 (posted twice)2 (posted twice)Duplicate webhook delivery
main111naive2 (posted twice)2 (posted twice)2 (posted twice)Worker killed mid-batch
main111naive000Two workers claim one residual
main111naive2 (posted twice)2 (posted twice)2 (posted twice)Replay of a consumed approval token
main111naive2 (posted twice)2 (posted twice)2 (posted twice)Lost response after a committed ledger write
main111naive2 (posted twice)2 (posted twice)2 (posted twice)Ledger returns an ambiguous 5xx
main100naive111
| Scenario | main | naive | ||||
|---|---|---|---|---|---|---|
| Key | Op id | None | Key | Op id | None | |
| Crash before commit | 1 | 1 | 1 | 2 (posted twice) | 2 (posted twice) | 2 (posted twice) |
| Duplicate webhook delivery | 1 | 1 | 1 | 2 (posted twice) | 2 (posted twice) | 2 (posted twice) |
| Worker killed mid-batch | 1 | 1 | 1 | 0 | 0 | 0 |
| Two workers claim one residual | 1 | 1 | 1 | 2 (posted twice) | 2 (posted twice) | 2 (posted twice) |
| Replay of a consumed approval token | 1 | 1 | 1 | 2 (posted twice) | 2 (posted twice) | 2 (posted twice) |
| Lost response after a committed ledger write | 1 | 1 | 1 | 2 (posted twice) | 2 (posted twice) | 2 (posted twice) |
| Ledger returns an ambiguous 5xx | 1 | 0 | 0 | 1 | 1 | 1 |
Three cells per branch, one per adapter capability: Adapter capability: Key = ENFORCES_KEY · Op id = BY_OPERATION_ID only · None = NONE / NONE
Adjustments posted at the simulated ledger for one unit of work, counted by the ledger itself. main never exceeds one; its two zeros are correct outcomes — resolved by query, or handed to an operator — not shortfalls.
Source: README · chaos suite results, generated by make chaos-table against real PostgreSQL
Architecture
- 01Input
Settlement file
ingest, or quarantine the batch
- 02Code
Deterministic match
per-currency tolerance bands
- 03Code
Exception + evidence pack
one per residual
- 04Model
Treatment code
closed enum · no numeric field · may abstain
- 05Human
Approval by role
controller approves · analyst may reject
- 06Code
Amount, account, period
pure Decimal calculator
- 07Store
Operation id + outbox
one transaction
- 08Effect
Dispatch
capability-declaring ledger adapter
Branch · Outcome UNKNOWN — by the adapter's declared capability · one of
- ↳01Effect
Re-send, same operation id
ENFORCES_KEY · bounded, inside the declared window
- ↳02Gate
Reconcile by query
BY_OPERATION_ID only
- ↳03Human
Manual recovery
neither · the automatic path stops
Everything outside the one model step is deterministic and unit-tested.
- InputArrives from outside the system
- CodeDeterministic code
- ModelA model proposal, confined to a closed output
- HumanA person decides
- GateDecides whether work proceeds
- StoreDurable state
- EffectAn irreversible or outbound effect
Engineering notes
The model's only channel is a closed enum
The model proposes a treatment — rebook, accrue, write_off or escalate — with a confidence band, citations and the option to abstain. Its response schema contains no numeric type anywhere in its tree, and a CI guard walks the schema and fails the build if one appears.
The rationale is provenance for humans only: no code parses it. The amount calculator takes the exception's persisted facts, a treatment code and system-owned ledger context, and a guard asserts it cannot import the proposal model. Where a treatment cannot be priced deterministically the answer is escalate — the correct label for 214 of the 250 golden records.
Five guarantees, deliberately kept apart
Conflating them is how the stronger claim gets asserted by accident, so the repository states each with its own condition.
- One claim per residual and one adjustment per operation id — unconditional.
- Transactional outbox: the intent is never lost — unconditional, and deliberately at-least-once.
- No second dispatch for a known terminal outcome — bounded by what the system can know.
- Adapters declare their capabilities and outcomes are three-valued — an adapter that cannot say UNKNOWN is rejected.
- An effectively-once financial side effect — only where the adapter enforces an idempotency key or exposes a queryable posting identity. Otherwise the claim is withdrawn, not reworded.
A kill test that could have failed
The reliability claim is proven against a deliberately naive branch that double-posts — a suite that passes on both branches proves nothing. Every scenario runs against three adapter capability configurations on real PostgreSQL, and every number is the simulated ledger's own applied count, never inferred from this system's records.
The results table in the README is generated from that run, and a check fails the build if it drifts. A five-mutant battery plants the ways such a gate can be green and worthless.
The live model measurement, and what contains the wrong answers
One bounded run over the 250-record golden set: 251 live calls against a declared ceiling of 750, 98.8% usable through the shipped path, p95 latency 8.15 s. The model was good at the judgement and bad at declining to make one: of the 211 escalate-labelled records it answered, 177 got a concrete treatment.
Three independent fail-closed controls contain different parts of that. During the run the citation check refused a hallucinated evidence id; the account policy refuses all 177, because no account is configured for those classifications and so no amount can be computed; and the approval gate stands in front of the single wrong answer that would have priced. Cost was not measured: the subscription-backed route returns no billing field, and no list-price estimate is made.
Screens



Limitations
As the project states them. Read these before relying on any number above.
- Every row is synthetic and the ledger is simulated; there is no real ledger integration.
- Nothing runs the stages in sequence as a service: the demo seeder composes them, ingestion is command-line driven, and the retry and reconciliation passes are bounded one-shot passes rather than daemons.
- The shipped package makes no live model call: the deployed demo uses a declared stand-in, and the one live measurement ran through a test-only transport on a workstation.
- The effectively-once effect is conditional on the ledger adapter's declared and proven capabilities.
- Under ENFORCES_KEY the suppression is performed by a simulated ledger written in the repository, so it shows the dispatcher behaving correctly given an enforcing ledger — not that any real ledger enforces anything.
- The free-tier backend sleeps; the first request after a quiet period can take up to a minute.
Facts and stack
- Golden set
- 250 records
- Chaos cells
- 7 scenarios × 3 adapter configurations × 2 branches
- Decision record
- 71 numbered ADRs, ADR-001 to ADR-071
- Demo sign-in
- one click, as a public analyst, operator or controller role
Stack
- Python 3.12
- FastAPI
- PostgreSQL 16
- SQLAlchemy 2
- Alembic
- Next.js 15
- TypeScript
- Docker
- GitHub Actions
- Vercel
- Render
- Neon
Skills shown
- Idempotency
- Transactional outbox
- Bounded retry & DLQ
- Deterministic money
- LLM integration
- Evaluation vs baselines
- Role-separated approval
- Audit trail