07Retrieval (RAG)
Parts Answer Gate
Free demo — may take ~1 min to wake.
In short
- What it does
- Answers technicians' questions from maintenance manuals, quoting the exact passage that applies to their machine on a given date.
- Why it matters
- An answer taken from an outdated manual or the wrong machine variant can look confident, cite a real document and still be wrong.
- What I built
- Built search that filters by machine, serial number and date before ranking, a gate that can refuse to answer, and word-for-word answers with citations.
- Key result
- It never returned an outdated or wrong-variant passage in 600 replayed queries (120 questions, each at 5 dates), but 4 of its 12 release criteria, fixed in advance, failed — a fifth passed only trivially — so it is published as a negative result.
Overview
A bitemporal RAG system for maintenance technicians: variant, serial and validity are filtered in SQL before anything is ranked. Four of its twelve pre-registered kill conditions failed, and the project is published as that.
- Problem
- A technician needs the procedure in force for their exact machine variant and serial on a given date. A retriever with no effectivity constraint can quote a superseded revision or another variant's value verbatim — a confident, well-cited, wrong answer — and filtering a ranked list afterwards silently answers from whatever survives.
- Built
- Hybrid BM25 and pgvector retrieval with variant, serial, validity and knowledge-time predicates in the SQL WHERE clause of both ranking queries; a gate of seven deterministic signals that answers, abstains or sends to review; extractive answers with verbatim citations at character offsets; in English, Turkish and Russian.
- Hard part
- Bitemporality. For example, a correction issued in 2025 about a 2021 procedure is valid in 2021 and known from 2025. Asked what the technician had in front of them at the time, the system returns the belief a later correction replaced — without rewriting it.
Evidence
release criteria, fixed before any source file existed, that failed — E, F, I and K; none was lowered, removed or disabled afterwards
A fifth, G, passes near-vacuously: for 297 of 312 hold-out questions the effectivity filter leaves the gate nothing to catch.
superseded passages returned across 600 date-specific queries (120 questions at 5 dates)
cited spans absent from the document they name
hold-out wrong-answer rate gated, against 0.0481 ungated
Both come from the 15 hold-out questions about a family the corpus does not contain — the one class where the effectivity filter cannot help. Gated: a single wrong answer. G is near-vacuous for this reason.
Pre-registered kill conditions
4 of 12 failedE · F · I · K
- APassNo superseded chunk for an as-of query
- BPassNo other variant's chunk
- CPassNo part number the evidence lacks
- DPassNo cited span absent from its document
- EFailThe gate refuses unsupportable questions
- FFailHold-out recall@10 ≥ 0.85 and above every baseline
- GPass · near-vacuousHold-out wrong-answer rate ≤ 0.02
- HPassUngated wrong answers ≥ 5× the gated rate
- IFailAbstention on the unanswerable set ≥ 0.90
- JPassTwo runs agree byte for byte
- KFailVector retrieval executes through pgvector
- LPassNo document appears in both splits
Thresholds fixed before any source file existed. G passes near-vacuously: for 297 of 312 hold-out questions the effectivity filter excludes, upstream of the gate, every passage that could make an answer wrong. The engineering is finished and deployed; the experiment's result is negative.
Source: README · generated from the graded kill test
Architecture
- 01Input
Question
as_of · known_as_of · variant · serial · language
- 02Code
Effectivity filter in SQL
WHERE clause of both ranking queries
- 03Code
Hybrid ranking
BM25 + pgvector
- 04Gate
Gate
seven deterministic signals · no model output
- 05Effect
Extractive answer
verbatim spans at character offsets
Branch · Gate withholds · one of
- ↳01Gate
ABSTAIN
carries no approved evidence
- ↳02Human
REVIEW
for example, two in-force sources disagree
A passage outside the asked-for variant, serial range or in-force window is never scored at all.
- InputArrives from outside the system
- CodeDeterministic code
- HumanA person decides
- GateDecides whether work proceeds
- EffectAn irreversible or outbound effect
Engineering notes
Why it failed, and one root cause under most of it
E and I are the gate over-covering: term coverage uses a bidirectional prefix match that errs towards covering. It refuses 44 of 45 questions about a product family that does not exist, but only 17 of 45 where the product exists and the attribute does not. F, G, H and K share a cause: the criteria were written as if the effectivity filter sat beside the thing being measured, when it sits upstream of everything.
The filter shrinks the candidate set to about 21 rows, for which PostgreSQL's planner uses the effectivity index rather than the vector index, so K cannot hold. On the hold-out the system's recall@10 of 0.9259 ties bm25_only and trails dense_only at 0.9352, so F fails. The failures stand as scored; the three corrections made after the first hold-out score are recorded in the decision log.
What worked, and is published
Effectivity filtered in SQL before ranking held in every replayed query: no superseded or wrong-variant passage over 600 queries (120 questions at 5 dates). The second temporal axis works: a historical knowledge query returns the belief a correction replaced. No ungrounded part number over 734 answers, no unfaithful span over 3,669 citations, and ten of ten planted breaches caught.
The live deployment is checked over HTTP by a committed script: seven cases, including one question at a single validity date that cites the original revision under earlier knowledge and its correction under current knowledge. The deployed database is read out of its own catalog: PostgreSQL 16.15, pgvector 0.8.0, and an HNSW index over 12,420 embedded chunk rows — 4,140 chunks in three languages.
Two benchmark iterations, both kept in git
The first benchmark was fully scored and then found not to be a valid retrieval test: after filtering, the median hold-out question left 9 eligible passages against top-k 10, so ranking could not change recall. It is recorded, unsquashed. The second iteration rebuilt the corpus on distinct content — 14 families, 31 variants, 84 documents, 4,140 chunks — leaving a median of 21 eligible passages and 0% at or below top-k.
Screens






Limitations
As the project states them. Read these before relying on any number above.
- Four of twelve pre-registered kill conditions fail (E, F, I, K); G passes near-vacuously.
- Three corrections were made after the second hold-out's first score. One re-froze it with 15 more unanswerable questions, which made G and H measurable and made I fail; the project asks readers to read that commit sceptically.
- No native speaker reviewed the corpus and no LLM judge was used: the multilingual result measures retrieval and gating over synthetic parallel text.
- The public instance serves query vectors computed at build time, because the multilingual encoder measured 671 MB resident against the free tier's 512 MB; free text outside the corpus is refused with an explanation.
- The managed-vector comparison used a local Qdrant container; Qdrant Cloud was never reached.
- The abstractive arm raises; no cost, latency or quality figure is published for any live model.
Facts and stack
- Corpus
- 14 families · 31 variants · 84 documents · 4,140 chunks
- Languages
- EN / TR / RU, reported separately
- Planted breaches
- 10 of 10 caught
- Encoder
- paraphrase-multilingual-MiniLM-L12-v2
Stack
- Python 3.12
- FastAPI
- PostgreSQL 16
- pgvector
- BM25
- fastembed
- Neon
- Docker
- Render
Skills shown
- RAG
- pgvector
- Hybrid retrieval
- Embeddings
- Citations
- Bitemporal data
- Pre-registered evaluation