Skip to content
Tahir Aslanli

07Retrieval (RAG)

Parts Answer Gate

Closed — pre-registered negative resultDeployed / live

Free demo — may take ~1 min to wake.

In short

What it does
Answers technicians' questions from maintenance manuals, quoting the exact passage that applies to their machine on a given date.
Why it matters
An answer taken from an outdated manual or the wrong machine variant can look confident, cite a real document and still be wrong.
What I built
Built search that filters by machine, serial number and date before ranking, a gate that can refuse to answer, and word-for-word answers with citations.
Key result
It never returned an outdated or wrong-variant passage in 600 replayed queries (120 questions, each at 5 dates), but 4 of its 12 release criteria, fixed in advance, failed — a fifth passed only trivially — so it is published as a negative result.

Overview

A bitemporal RAG system for maintenance technicians: variant, serial and validity are filtered in SQL before anything is ranked. Four of its twelve pre-registered kill conditions failed, and the project is published as that.

Problem
A technician needs the procedure in force for their exact machine variant and serial on a given date. A retriever with no effectivity constraint can quote a superseded revision or another variant's value verbatim — a confident, well-cited, wrong answer — and filtering a ranked list afterwards silently answers from whatever survives.
Built
Hybrid BM25 and pgvector retrieval with variant, serial, validity and knowledge-time predicates in the SQL WHERE clause of both ranking queries; a gate of seven deterministic signals that answers, abstains or sends to review; extractive answers with verbatim citations at character offsets; in English, Turkish and Russian.
Hard part
Bitemporality. For example, a correction issued in 2025 about a 2021 procedure is valid in 2021 and known from 2025. Asked what the technician had in front of them at the time, the system returns the belief a later correction replaced — without rewriting it.

Evidence

4 of 12Negative result

release criteria, fixed before any source file existed, that failed — E, F, I and K; none was lowered, removed or disabled afterwards

A fifth, G, passes near-vacuously: for 297 of 312 hold-out questions the effectivity filter leaves the gate nothing to catch.

0

superseded passages returned across 600 date-specific queries (120 questions at 5 dates)

0 / 3,669

cited spans absent from the document they name

0.0047

hold-out wrong-answer rate gated, against 0.0481 ungated

Both come from the 15 hold-out questions about a family the corpus does not contain — the one class where the effectivity filter cannot help. Gated: a single wrong answer. G is near-vacuous for this reason.

Pre-registered kill conditions

4 of 12 failedE · F · I · K

  1. APassNo superseded chunk for an as-of query
  2. BPassNo other variant's chunk
  3. CPassNo part number the evidence lacks
  4. DPassNo cited span absent from its document
  5. EFailThe gate refuses unsupportable questions
  6. FFailHold-out recall@10 ≥ 0.85 and above every baseline
  7. GPass · near-vacuousHold-out wrong-answer rate ≤ 0.02
  8. HPassUngated wrong answers ≥ 5× the gated rate
  9. IFailAbstention on the unanswerable set ≥ 0.90
  10. JPassTwo runs agree byte for byte
  11. KFailVector retrieval executes through pgvector
  12. LPassNo document appears in both splits

Thresholds fixed before any source file existed. G passes near-vacuously: for 297 of 312 hold-out questions the effectivity filter excludes, upstream of the gate, every passage that could make an answer wrong. The engineering is finished and deployed; the experiment's result is negative.

Source: README · generated from the graded kill test

Architecture

  1. 01Input

    Question

    as_of · known_as_of · variant · serial · language

  2. 02Code

    Effectivity filter in SQL

    WHERE clause of both ranking queries

  3. 03Code

    Hybrid ranking

    BM25 + pgvector

  4. 04Gate

    Gate

    seven deterministic signals · no model output

  5. 05Effect

    Extractive answer

    verbatim spans at character offsets

Branch · Gate withholds · one of

  • ↳01Gate

    ABSTAIN

    carries no approved evidence

  • ↳02Human

    REVIEW

    for example, two in-force sources disagree

A passage outside the asked-for variant, serial range or in-force window is never scored at all.

  • InputArrives from outside the system
  • CodeDeterministic code
  • HumanA person decides
  • GateDecides whether work proceeds
  • EffectAn irreversible or outbound effect

Engineering notes

Why it failed, and one root cause under most of it

E and I are the gate over-covering: term coverage uses a bidirectional prefix match that errs towards covering. It refuses 44 of 45 questions about a product family that does not exist, but only 17 of 45 where the product exists and the attribute does not. F, G, H and K share a cause: the criteria were written as if the effectivity filter sat beside the thing being measured, when it sits upstream of everything.

The filter shrinks the candidate set to about 21 rows, for which PostgreSQL's planner uses the effectivity index rather than the vector index, so K cannot hold. On the hold-out the system's recall@10 of 0.9259 ties bm25_only and trails dense_only at 0.9352, so F fails. The failures stand as scored; the three corrections made after the first hold-out score are recorded in the decision log.

What worked, and is published

Effectivity filtered in SQL before ranking held in every replayed query: no superseded or wrong-variant passage over 600 queries (120 questions at 5 dates). The second temporal axis works: a historical knowledge query returns the belief a correction replaced. No ungrounded part number over 734 answers, no unfaithful span over 3,669 citations, and ten of ten planted breaches caught.

The live deployment is checked over HTTP by a committed script: seven cases, including one question at a single validity date that cites the original revision under earlier knowledge and its correction under current knowledge. The deployed database is read out of its own catalog: PostgreSQL 16.15, pgvector 0.8.0, and an HNSW index over 12,420 embedded chunk rows — 4,140 chunks in three languages.

Two benchmark iterations, both kept in git

The first benchmark was fully scored and then found not to be a valid retrieval test: after filtering, the median hold-out question left 9 eligible passages against top-k 10, so ranking could not change recall. It is recorded, unsquashed. The second iteration rebuilt the corpus on distinct content — 14 families, 31 variants, 84 documents, 4,140 chunks — leaving a median of 21 eligible passages and 0% at or below top-k.

Screens

An answered question: the extracted span and the first citation row, under the release-gate failure banner
An answered question: the extracted span and the first of its citations, under the release-gate failure banner that heads every page.Captured from the live deployment
A question about the AX7-165 at 29 November 2020 under current knowledge, citing the correction document
Valid at 29 November 2020, under current knowledge: the citations come from the correction (…-manc-b-en).Captured from the live deployment
The same question and validity date, known at 29 November 2020, citing the original revision
The same question and validity date, known at 29 November 2020: the citations come from the original revision (…-man-b-en).Captured from the live deployment
Kill condition E failing on the live deployment: a Turkish question the corpus cannot support, answered anyway
Kill condition E failing, live: the corpus has no passage that answers this, and the gate answers anyway. Published, not cropped out.Captured from the live deployment
A question the corpus cannot support, withheld by the gate
A question the corpus cannot support, withheld.Captured from the live deployment
Top of the evidence page: hold-out retrieval figures and the start of the baseline table, under the failure banner
The top of the evidence page: hold-out figures and the start of the baseline table, under the failure banner.Captured from the live deployment

Limitations

As the project states them. Read these before relying on any number above.

  • Four of twelve pre-registered kill conditions fail (E, F, I, K); G passes near-vacuously.
  • Three corrections were made after the second hold-out's first score. One re-froze it with 15 more unanswerable questions, which made G and H measurable and made I fail; the project asks readers to read that commit sceptically.
  • No native speaker reviewed the corpus and no LLM judge was used: the multilingual result measures retrieval and gating over synthetic parallel text.
  • The public instance serves query vectors computed at build time, because the multilingual encoder measured 671 MB resident against the free tier's 512 MB; free text outside the corpus is refused with an explanation.
  • The managed-vector comparison used a local Qdrant container; Qdrant Cloud was never reached.
  • The abstractive arm raises; no cost, latency or quality figure is published for any live model.

Facts and stack

Corpus
14 families · 31 variants · 84 documents · 4,140 chunks
Languages
EN / TR / RU, reported separately
Planted breaches
10 of 10 caught
Encoder
paraphrase-multilingual-MiniLM-L12-v2

Stack

  • Python 3.12
  • FastAPI
  • PostgreSQL 16
  • pgvector
  • BM25
  • fastembed
  • Neon
  • Docker
  • Render

Skills shown

  • RAG
  • pgvector
  • Hybrid retrieval
  • Embeddings
  • Citations
  • Bitemporal data
  • Pre-registered evaluation