MILLENNIUM RESEARCH Faithfulness Eval Suite The Suite Inspector Measured Benchmark Audit Deliverables Pricing Request a scored pilot
LaTeX → Lean 4 · a private faithfulness eval, built from real failures

We measure whether formal statements say what the paper says, with error bars we publish on ourselves.

{{ heroA }}
Human verdicts
every certification is a person, never a model
{{ heroB }}
Certified faithful pairs
theorems and definitions · novelty-gated
{{ heroC }}
Attributed real failures
produced by a real system, caught by a human
{{ heroD }}
Frontier escape rate
human-rejected pairs that pass both frontier judges
139 / 139
Counterexample record
every counterexample-backed flag upheld by human review, across seven audits
91.8%
Verified flag precision
156 of 170 sampled flags survived human review, pooled across audits
Request a scored pilot See the audits
All figures measured 2026-07-19 · reject-only invariant: agents may reject, only humans certify
The Suite

Five sections, and the asset never ships whole

You submit outputs; we return scored reports. The splits are stratified by difficulty and failure class, so seeing a sample reveals nothing about the private core.

#
Section
Items
Visibility
Purpose
Status
{{ row.rank }}
{{ row.name }}
{{ row.sub }}
{{ row.n }}
{{ row.vis }}
{{ row.purpose }}
{{ row.status }}
gold rule: no item is sold or scored without 3 blind human labels

Every negative is a real failure a production autoformalizer emitted and a human rejected, not a programmatic perturbation. CC0/CC-BY sources only; no proof text ships anywhere; the automated leak audit runs before every export and its history is committed to git.

Example Inspector

Straight from the corpus

Every example below is verbatim from the corpus: the source paper, the Lean draft, the frontier judge’s back-translation, and the human verdict of record. The flagged ones are from the consensus-breaking section: both frontier judges passed them and a human caught them.

{{ pairsMeta }}
{{ p.id }} {{ p.kind }}
{{ p.tag }}
{{ selId }} {{ selKind }}
{{ selVLabel }} {{ selCatLabel }}
{{ selDiag }}
Detection ladder
{{ l.name }} {{ l.out }}
recorded outcomes for this exact pair, replayed from the corpus evidence; nothing here is illustrative
Source statement · LaTeX
{{ selLatex }}
Formal statement · Lean 4
{{ selLean }}
Provenance trail
{{ accRIcon }} Frontier judge: verbatim back-translation judge A · v1 prompt
“{{ selReason }}”
{{ accPIcon }} Strict two-judge consensus both must pass
{{ selConsensus }}
{{ accHIcon }} Human verdict of record reject-only invariant
{{ selReviewer }}
Measured, Not Promised

The numbers we publish on ourselves

Every eval sells trust. Ours ships its own error bars: the judge configuration behind our harness was calibrated against 886 frozen human verdicts.

LLM judges miss what humans catch
23.4% escape rate
Theorems17.0% Definitions35.6%

23.4% of human-rejected pairs pass both frontier judges under strict consensus. An eval gold-labeled by an LLM judge inherits exactly this blind spot. That is why our gold tier is human, triple-reviewed, and blind.

The corpus tracks a real system getting better
Hard-cheat rate, first drafting era21%
Hard-cheat rate, current era3%
measured on human verdicts, era over era

The negatives span every era, so the suite contains both the crude cheats early systems make and the subtle drift strong systems still make: the failure distribution a buyer faces.

When our screen rejects, it’s right
99.8%rejection precision

Every one of 446 consecutive automated rejections from a full production batch was hand-audited: 445 correct and 1 false rejection, which we found and reopened. The audit itself is part of the product discipline.

Known limits, stated plainly
Our in-house negatives originate from one production system; external coverage now spans three public benchmarks and three open autoformalizers (audits below).
Blind multi-rater gold labeling is in progress; inter-rater κ ships only once measured.
Consensus-section class shares are approximate at n=59.
Public Benchmark Audit

We ran our screen over the field’s benchmarks

A full census: every statement, no sampling. A flag is a candidate, not a verdict: each carries judge back-translations and, where the probe found one, a concrete counterexample, and the verified column fills only from human spot-check. These corpora circulated for years; the defects were sitting in them the whole time, including two nobody had reported and ten introduced by the expert corrections themselves.

Benchmark
Pairs
Flagged
Counterex.
Verified precision
Notes
{{ b.name }}
{{ b.sub }}
{{ b.n }}
{{ b.flagged }}
{{ b.ce }}
{{ b.pill }}
{{ b.notes }}
measured 2026-07-16 · dual-judge strict consensus + falsification probe + numeric cross-evaluation · Lean elaboration (our mathlib pin): 99.2% / 99.6% / 94.6% validflagged ≠ wrong; passed_screening ≠ certified either; humans decide both directions
Paired control: the flags track the fixes

Same detector, same day, over miniF2F v1 and over the expert-corrected v2s. On exactly the statements the experts changed, our v1 flags cleared. On the statements they left alone, they didn’t.

Flags cleared where experts corrected the statement63.9% (46/72)
Flags cleared where they didn’t11.1% (1/9)
paired control · n=487 aligned statements

Honest floor: judges are stochastic, so unchanged statements still flip flags between any two runs (5.3% → 11.8% here). That is why we publish the differential, never the raw drop.

Counterexamples on the corrected benchmark
20 / 20counterexample flags land on expert-corrected statements

All 20 were confirmed by hand. Each carried a concrete counterexample and sat on a statement the experts had touched: errors that survived expert correction, caught by the screen and verified by a human.

Then we ran it over the field’s autoformalizers

Four open autoformalizers, the same 488 miniF2F problems, one draft per system under each model’s own documented prompt and decoding settings, scored by the identical screen as the audit above. The point is not a ranking: a large share of translations Lean accepts do not say what the problem says, and a compile check catches none of it. Every defect below carries evidence a human confirmed.

System
Problems
Flagged
Counterex.
Verified precision
Notes
{{ b.name }}
{{ b.sub }}
{{ b.n }}
{{ b.flagged }}
{{ b.ce }}
{{ b.pill }}
{{ b.notes }}
measured 2026-07-19/20 · identical protocol to the benchmark audit · verbatim upstream prompts, params in a committed manifest · elaborates in our mathlib pin: 81.3% / 55.9% / 88.1% / 97.3% (environment-sensitive; not counted as defects)
139 / 139counterexample-backed flags human-confirmed, across every audit we have run: three public benchmarks and four autoformalizers

When this screen produces a concrete counterexample, values where the hypotheses hold and the conclusion fails, human review has upheld it every time to date. That class is the backbone of every number we publish.

Deliverables

What an engagement ships

Submit your pairs; your data never enters our database and every run carries a hard spend cap.

Scored audit reportMD / PDF
Flag rates with Wilson intervals, split by kind and failure class
Per-pair evidence: judge back-translations, lint findings, counterexamples
A separate consensus-breaking score line: beating one LLM judge vs beating the measured frontier
Our own error bars printed in every report’s calibration block
Scoring harnessLive
$ score-external your_pairs.jsonl --budget 50 --probe
20 pairs; est. ~$1.60 · credits: OK
pair-00014 -> flagged (deterministic-vacuous)
pair-00011 -> flagged (judge-a + judge-b mismatch)
pair-00002 -> passed_screening
evidence: scores.jsonl · report: scores.md
passed_screening ≠ certified; humans certify

Deterministic screens run free before any paid judge call; a crash or budget breach loses nothing (per-pair appends, resume by id).

Consensus-breaking sectionHard
59pairs both frontier judges passed; a human rejected (28 theorems, 31 definitions)

A further 19 candidates carry external provenance: independent benchmark experts corrected statements our judges had passed.

Missing context Plausible neighbor Fragment capture Type semantics Transpositions Alias defs
Three classes have shipped mechanical detectors (≤2% fire rate on faithful pairs). The same detector stack validated out-of-house: 139 of 139 counterexample-backed flags across three public benchmarks and four autoformalizers were human-confirmed; see the audits above.
Pricing

Priced against the market, in the open

Most eval vendors quote nothing until a sales call. We publish our entry price and the market anchors we derived it from, so you can check our math before you send us yours.

Scored pilotStart here
$5,000
flat fee, two-week turnaround
Up to 1,000 of your pairs through the full detector stack
Per-pair evidence file and the scored audit report, calibration block included
A 30-pair slice human-verified, so you can audit the audit
Request a pilot
Suite subscription
from$15,000
per quarter, scoped to cadence and volume
Recurring scored runs against the private suite, consensus score line included
Sealed refresh reserve rotates in as disclosed items retire
Semi-private section access under NDA for your own spot-checks
the pilot fee credits toward a first subscription quarter
Human certification and commissioned items
from$40
per pair certified, volume-scoped
A human verdict of record on any slice of your corpus, reject-only invariant intact
Commissioned consensus-grade items, $500 to $2,500 each, scoped by difficulty
Custom failure-class sections built from your own system’s output

market anchors behind these numbers: expert mathematician time runs $110 to $130 per hour (Mercor); commissioned benchmark problems run $300 to $1,000 (FrontierMath) with top-tier items to $5,000 (Humanity’s Last Exam); commodity LLM-judge calls run $10 to $20 per thousand (Patronus), and our calibration section shows why those alone are not enough.

Millennium Research · millenniumresearch.ai · All figures measured on our production corpus, 2026-07-19 · methods documented in eval-suite-methods.md