You submit outputs; we return scored reports. The splits are stratified by difficulty and failure class, so seeing a sample reveals nothing about the private core.
Every negative is a real failure a production autoformalizer emitted and a human rejected, not a programmatic perturbation. CC0/CC-BY sources only; no proof text ships anywhere; the automated leak audit runs before every export and its history is committed to git.
Every example below is verbatim from the corpus: the source paper, the Lean draft, the frontier judge’s back-translation, and the human verdict of record. The flagged ones are from the consensus-breaking section: both frontier judges passed them and a human caught them.
Every eval sells trust. Ours ships its own error bars: the judge configuration behind our harness was calibrated against 886 frozen human verdicts.
23.4% of human-rejected pairs pass both frontier judges under strict consensus. An eval gold-labeled by an LLM judge inherits exactly this blind spot. That is why our gold tier is human, triple-reviewed, and blind.
The negatives span every era, so the suite contains both the crude cheats early systems make and the subtle drift strong systems still make: the failure distribution a buyer faces.
Every one of 446 consecutive automated rejections from a full production batch was hand-audited: 445 correct and 1 false rejection, which we found and reopened. The audit itself is part of the product discipline.
A full census: every statement, no sampling. A flag is a candidate, not a verdict: each carries judge back-translations and, where the probe found one, a concrete counterexample, and the verified column fills only from human spot-check. These corpora circulated for years; the defects were sitting in them the whole time, including two nobody had reported and ten introduced by the expert corrections themselves.
Same detector, same day, over miniF2F v1 and over the expert-corrected v2s. On exactly the statements the experts changed, our v1 flags cleared. On the statements they left alone, they didn’t.
Honest floor: judges are stochastic, so unchanged statements still flip flags between any two runs (5.3% → 11.8% here). That is why we publish the differential, never the raw drop.
All 20 were confirmed by hand. Each carried a concrete counterexample and sat on a statement the experts had touched: errors that survived expert correction, caught by the screen and verified by a human.
Four open autoformalizers, the same 488 miniF2F problems, one draft per system under each model’s own documented prompt and decoding settings, scored by the identical screen as the audit above. The point is not a ranking: a large share of translations Lean accepts do not say what the problem says, and a compile check catches none of it. Every defect below carries evidence a human confirmed.
When this screen produces a concrete counterexample, values where the hypotheses hold and the conclusion fails, human review has upheld it every time to date. That class is the backbone of every number we publish.
Submit your pairs; your data never enters our database and every run carries a hard spend cap.
Deterministic screens run free before any paid judge call; a crash or budget breach loses nothing (per-pair appends, resume by id).
A further 19 candidates carry external provenance: independent benchmark experts corrected statements our judges had passed.
Most eval vendors quote nothing until a sales call. We publish our entry price and the market anchors we derived it from, so you can check our math before you send us yours.
market anchors behind these numbers: expert mathematician time runs $110 to $130 per hour (Mercor); commissioned benchmark problems run $300 to $1,000 (FrontierMath) with top-tier items to $5,000 (Humanity’s Last Exam); commodity LLM-judge calls run $10 to $20 per thousand (Patronus), and our calibration section shows why those alone are not enough.