This is the nerd page. Everything the front page summarizes is derived from the record below: how the engine works, how it's measured, every number with its raw run attached, and the exact commands to reproduce all of it on your own machine. If a claim on this site can't be traced to a file in the repository, treat it as false and tell us.
v1.0 record — 118-document battery + the utility axis, measured 2026-07-19 → 2026-07-24, engine cb3627c, published as drawnTwo independent detection layers over one deterministic masking core. The design rule throughout: recall lives in code; the model only classifies and nominates under verbatim verification. Anything the model says that can't be found verbatim in the document is refused, not guessed at.
A number without its denominator and its sampling story is marketing. Here is ours, stated before the results so it can't be quietly adjusted after them.
Any core-identity occurrence present in the exported text. No partial credit: a surname surviving inside a compound, a glued print artifact, a shorthand in a footer — all leaks. The gate's floor is zero core leaks; one leak exits red.
Every number links to its raw record in the repository. Nothing is blended: in-sample and out-of-sample are separate rows, and any tier whose misses a fix set ever drew on carries its saturation label — that is what the label column is for.
cb3627c, 2026-07-21)| Tier | Documents | Entities | Recall | Sample status |
|---|---|---|---|---|
| corpus — the ratchet floor | 8 | 243/243 | 100.00% | in-sample, six generations unbroken |
| long-form 79k words | 2 | 260/262 | 99.24% | in-sample |
| walk-forward round 1 | 5 | 78/78 | 100.00% | rails-saturated |
| walk-forward round 2 | 5 | 258/260 | 99.23% | mostly saturated |
| walk-forward round 3 | 36 | 610/667 | 91.45% | partially saturated — fix set 4 drew on its misses |
| walk-forward round 4 | 44 | 1,340/1,394 | 96.13% | partially saturated |
| round 5 — the headline | 18 | 698/735 | 94.97% | pure out-of-sample, pre-registered publish-as-drawn |
| battery total | 118 | 3,487/3,639 | 95.82% | — |
Round 5 by occurrences: 4,018/4,093 printed occurrences masked (98.2%) · 9/11 shorthand. The walk-forward curve, final for v1.0: 84.62 → 95.38 → 88.46 → 93.76 → 94.97%. Pre-registration (FINDINGS.md, recorded 2026-07-21 before any round-5 document existed): "v1.0 publishes round 5's number WHATEVER IT IS" — honored. Junk is counted raw, not yet as a rate: 47 flagged suspects across the internal ratchet, 0–9 table rows per document on round 5 (28 total); a denominator-honest junk rate is queued and stays unclaimed until it is measured.
Per-class rollups exist for the two exhaustively-annotated in-sample corpora. No per-class aggregate has been computed for the out-of-sample rounds yet — round-5 misses carry class tags individually (see the leak ledger below), and an out-of-sample per-class rollup is queued; until it lands, no per-class number here should be read as an out-of-sample claim.
| Identity class | hardening matrix — 33 docs, in-sample | 79k long-form — in-sample, after fixes |
|---|---|---|
| Persons | 143/143 | 35/35 |
| Companies & organizations | 190/190 | 160/162 |
| Brands & marks | 68/68 | 24/25 |
| IDs & accounts | 75/75 | 21/22 |
| Addresses | 64/64 | 15/15 |
| Phones | 49/49 | 3/3 |
| Emails | 30/30 | — |
An entity counts only when every printed variant of it has zero residual in the export — the strictest scoring direction, applied identically to every system below.
| System | Scope | Result |
|---|---|---|
| Microsoft Presidio, as shipped | 90 walk-forward documents, our ground truth | 1,034/2,399 — 43.10% (per round: 51.3 · 33.5 · 33.9 · 48.9%) |
| Our engine, same 90 documents | rounds 1–4, same metric | 84.62 · 95.38 · 88.46 · 93.76% |
Everything above measures what the export withholds. This measures what it still carries — the number that matters when the redacted copy's whole purpose is to be analyzed by a frontier AI that was never allowed to see the original. Protocol (2026-07-24, harness committed): for each round-5 document a question-writer read the original and drafted exactly 5 analytical questions — obligations, procedure, outcomes, remedies, risk, findings; any question answerable by a name, date, address or other identity is banned, and because the writer saw only the original, questions cannot skew toward what a masked copy answers well. Two blind analysts then answered independently — one from the original, one from only the masked export, placeholders declared legitimate actors — and a judge scored each pair for substance equivalence. utility = (equivalent + ½·partial) / questions. The masked side is the round-5 export byte-for-byte as scored above — the same files that measured 94.97% / 98.2%.
| Measure | Result |
|---|---|
| Aggregate — 90 questions across all 18 documents | 0.983 — 87 equivalent · 3 partial · 0 divergent |
| Documents where every answer matched (5/5 equivalent) | 15 of 18 |
| Questions where the masked reader contradicted the original | 0 |
Same rule as the leak ledger: every non-equivalent verdict publishes, mechanism named. govdoc-cfpb — the masked reader could not name the bankruptcy forum or docket; the forum is an identity and was masked by design — the question brushed the boundary the metric excludes. ocr-mapleton1965 — the net-to-Surplus figure and two column totals were unreadable: an over-mask (three summary dollar figures caught by an ID rail) — exactly the class of row the mandatory review exists to unmask before export. transcript-r5-3 — with the company name masked, the analyst tied a divestiture rationale to the wrong business line; industry context rides in names. Judge reasons verbatim, all 90 question/answer/verdict transcripts, and the harness as run are committed: RESULT.md · v1-round5-result.json · harness-v1.js. Analysts and judge: Claude Opus at high reasoning effort, 72 agent runs, zero errors. Honest bounds: five questions per document sample its substance, they do not exhaust it; the judge is the same model family as the analysts; identity questions are excluded by design — a redacted export deliberately cannot answer “who”, and that is the product working, not a loss. The Singapore document's answer transcripts are pruned from the committed artifact under the same redistribution ruling as its source text; its verdicts (5/5 equivalent) are retained and the recount is unaffected.
Every missed entity of the headline round, verbatim as recorded in round5-results.json: 37 core + 2 shorthand. Family labels come from the round's miss taxonomy (RESULT.md); rows marked — carry no taxonomized family. Diagnosis and fixes belong to fix set 5, provable only on round 6's fresh documents. This table is the reason to trust the ones above.
| Missed span | Document | Class | Family |
|---|---|---|---|
| Amendment No. 1 | employment-r5-1 | — | — |
| 001-35231 | employment-r5-2 | — | — |
| 10.2 | employment-r5-2 | — | — |
| 10.3 | employment-r5-2 | — | — |
| 104 | employment-r5-2 | — | — |
| Amendment No. 1 (as printed) | merger-r5-1 | — | — |
| AMENDMENT NO. 2 TO AGREEMENT AND PLAN OF MERGER | merger-r5-2 | — | — |
| Bancrofts | ocr-goshen1961 | COMPANY | — |
| Science Research Associates | ocr-goshen1961 | COMPANY | — |
| MELVIN G. HIGGINS | ocr-mapleton1965 | — | — |
| Int. Harvester Co. Int. Truck Rep. | ocr-mapleton1965 | — | — |
| 41022 | ocr-mapleton1965 | — | — |
| 2-7371 | ocr-mapleton1965 | — | — |
| 2-2011 | ocr-mapleton1965 | — | — |
| 974.102 | ocr-mapleton1965 | — | — |
| the assessors | ocr-mapleton1965 | SH | — |
| Dorothy Ballan-tyne | opinion-defamation2 | PERSON | OCR hyphen-wrapped person print |
| the Board’s | opinion-defamation2 | SH | — |
| Fuse | proxy-r5-1 | — | — |
| Opening Act | proxy-r5-1 | — | — |
| S&P 500 Index (as printed) | proxy-r5-1 | — | — |
| Digital Next | proxy-r5-1 | — | — |
| Device Care Center | proxy-r5-1 | — | — |
| HC/SUM 474/2024 | sg-judgment-r5 | ID | SG form code — rail mechanism bug, queued |
| HC/RA 141/2024 | sg-judgment-r5 | ID | SG form code — rail mechanism bug, queued |
| AD/OA 17/2025 | sg-judgment-r5 | ID | SG form code — rail mechanism bug, queued |
| CA/OA 17/2025 | sg-judgment-r5 | ID | SG form code — rail mechanism bug, queued |
| LC carbine | transcript-r5-1 | BRAND | — |
| 1022 Rifle | transcript-r5-1 | BRAND | — |
| Michael Roxland | transcript-r5-3 | PERSON | transcript bare-surname person |
| Greif Business System 2.0 | transcript-r5-3 | BRAND | — |
| Fiscal First Quarter 2025 Earnings Results Conference Call | transcript-r5-3 | ID | event-title ID |
| Austell, Georgia | transcript-r5-3 | ADDRESS | — |
| Fitchburg, Massachusetts | transcript-r5-3 | ADDRESS | — |
| Alexander Waters | transcript-r5-4 | PERSON | transcript bare-surname person |
| Steven Fisher | transcript-r5-4 | PERSON | transcript bare-surname person |
| Uniti | transcript-r5-4 | COMPANY | bare company short-form |
| Gigapower | transcript-r5-4 | COMPANY | bare company short-form |
| DY | transcript-r5-4 | ID | bare ticker |
Prior measurement campaigns, oldest first, each with its full record in the tree. History is append-only: superseded numbers stay published with what superseded them.
| Date | Campaign | Record |
|---|---|---|
| 2026-07-19 | Founding bench — layer-by-layer measurement, chunk-geometry sweeps, the carry-list lesson | EVIDENCE_2026-07-19.md |
| 2026-07-19 | Hardening profile — corpus growth, stress suite, leak taxonomy, the two-engine configuration | hardening-2026-07-19/REPORT.md |
| 2026-07-20 | Scale ladder — two real annual reports to 79k words, before/after published unedited | bench/scale/SCALE.md |
| 2026-07-20 → 21 | Walk-forward programme, rounds 1–5 — the generalization record and this page's headline | bench/oos-walkforward/ |
| 2026-07-21 | v1.0 release battery — the full 118 documents, engine cb3627c, wall-clock 12:23–21:22 |
bench/RELEASE_BATTERY_v1.md |
| 2026-07-24 | Utility benchmark v1 — the analysis-survives axis: 90 blind original-vs-masked questions over round 5, judged for substance equivalence | bench/utility/RESULT.md |
Every number above is attached to an exact, hash-verifiable configuration. If your hashes match, you are running what we measured.
The model is vanilla by construction and by check:
scripts/verify-model.mjs hashes your local weights against the pin. Adaptation
lives entirely in prompts-as-frozen-contracts and code rails — never a fine-tune, so the
weights stay auditable against Google's published artifact.
The corpus, ground truth, runner and floor ship in the tree. Reproduction is two commands after clone; the bench exits red on a single core leak.
The live bench needs the model running locally (see bench/BENCH.md for the recipe). Found a leak we didn't? That's a bug report, and it becomes a pinned test — the ratchet only turns one way.
Kept in the repo's constitution, enforced in review. If you catch this site making one of these, that's a bug: