Evidence snapshot · 1 October 2026

Benchmark and measurement notes

This snapshot reports local checks against generated and licensed fixture truth, plus a separate comparison with LibreOffice numbering output. It does not compare Agmt Verbatim with other parsers, and it does not measure production Worker performance.

These results support specific fixture checks only. They do not establish a competitor ranking or a general preservation guarantee. LibreOffice is not Microsoft Word; Word parity for numbering is unverified.

Download aggregate results (JSON)

Current evidence

Local fixture suite

306 / 306

Checked fixtures passed in the full local acceptance report.

Authored text ranges

2,532 / 2,532

Generated, mutated, and phase 2 fixtures across the checked views.

Original text ranges

360 / 360

Authored text ranges across 20 originals; independent view checks pass.

LibreOffice numbering comparison

98.89%

3,295 / 3,332 labels match in each view, counting unaligned labels as non-matches.

Local acceptance evidence and section 2 targets
Measure Target Current evidence Scope
Exact text, accepted / rejected / marked views 100% synthetic and real 2,532 / 2,532 synthetic and mutated; 360 / 360 original authored ranges Local authored-truth checks
Revision preservation 100% synthetic; ≥99.5% real 1,002 / 1,002 eligible checks; original revision truth not measured Local fixtures; no reviewed original revision annotations
Comment and anchor preservation 100% synthetic; ≥99% real 40 / 40 eligible checks; original comment truth not measured Local fixtures; no reviewed original comment annotations
Numbering in generated hard set 100% in both views Accepted 66 / 66; rejected 67 / 67 Local fixture truth
Numbering compared with LibreOffice ≥98% real target in each view Accepted and rejected: 3,295 / 3,332 exact (98.89%); 3,303 / 3,332 aligned (99.13%) 20 originals; LibreOffice output, not Word output
Defined-term F1 ≥0.98 synthetic; ≥0.95 real Synthetic accepted and rejected: 48 / 48 (1.00); real not measured Real legal annotations await independent review
Field cross-references 100% in both views Synthetic accepted 5 / 5; rejected 5 / 5; real not measured Local authored truth; real legal review pending
Text cross-reference precision / recall ≥0.98 / ≥0.95 synthetic; ≥0.95 / ≥0.90 real Synthetic accepted and rejected: 100% / 100%; real not measured Local authored truth; real legal review pending
100-page p95 CPU ≤1.5 s Not measured on the implemented production Worker Phase 1 prototype measurement is separate
100-page peak memory ≤64 MB Not measured on the implemented production Worker Phase 1 sampled allocation envelope is not a production peak
5 MB p95 end-to-end latency ≤3 s Not measured on the public Worker Requires public Worker measurement
MCP Registry and Claude directory adoption Listed within 12 weeks Not measured; no production listing Production is not live
Weekly keyless documents ≥100 within 12 weeks Not measured Production is not live
Integrating legal-AI builders ≥2 within 12 weeks Not measured Production is not live
Comparison with Pandoc, MarkItDown, Docling, and Adeu Same public subset and published method Pending; no baseline results are published No comparative ranking is available

Separate prototype resource observation

A Phase 1 throwaway parser was measured locally with a 100-page fixture. Its 20-run process CPU p95 was 160 ms and its sampled allocation envelope was 76,754,072 bytes (about 76.75 MB), at reported confidence 0.9. This was a feasibility observation, not the implemented engine or production Worker result.

The sampled envelope is not a proved absolute peak. Local workerd does not enforce Cloudflare's production isolate limit; the current 64 MB production memory target remains unmeasured.

Methodology and corpus

  1. Local correctness metrics are computed from generated fixtures, scripted mutations, phase 2 fixtures, and authored text ledgers. The full acceptance report covers 306 fixture cases; original-document text checks cover 20 source files. These results describe checked local fixtures, not arbitrary uploaded agreements.
  2. The separate numbering comparison renders accepted and rejected views through LibreOffice and compares labels with engine output. It reports both the whole set and source-alignment coverage. Unaligned labels count as non-matches in the whole-set rate. LibreOffice is an independent oracle, not Microsoft Word.
  3. The Phase 1 resource observation used a throwaway parser in local workerd with sampled allocation measurement. It has explicit measurement limitations and does not establish production Worker CPU, peak memory, or network latency.
  4. The planned external comparison against Pandoc, MarkItDown, Docling, and Adeu has not been run. Tool versions, outputs, and comparative methodology are therefore not reported here.

The checked corpus includes 20 original agreements. Its source ledger records provenance and license metadata. Read the license terms: Creative Commons Attribution 4.0 and Open Government Licence v3.0. Review of redistribution scope for some source material remains open; no endorsement by a publisher or government is implied.

Publisher source collections: Common Paper standards and UK government mutualisation templates.

Downloads and limits

Aggregate metric summary (JSON)

The source-bound raw oracle reports are not linked here. The numbering report contains document-derived labels and unmatched source text; this page and its downloadable summary expose aggregate counts only.

No report here establishes a general preservation rate, a competitor advantage, production performance, or a guarantee about Microsoft Word rendering.