What we measure, and what we found.

Benchmark audits

A full census of miniF2F, the expert-corrected miniF2F v2, and ProofNet#: 1,347 statements, every one screened, every flag verified by hand before anything went public. 76 confirmed defects, documented item by item, all three filings public. See the miniF2F filing, the v2 issue, and the ProofNet# filing.

Cross-system audits

Four open autoformalizers ran the same 488 problems under their own documented prompts and decoding settings. Goedel V2 failed the faithfulness screen least (21.7%), StepFun and Kimina near a third (32.6% and 34.0%), Herald’s single-draw outputs most (55.5%). All 80 sampled flags survived human review. A compile check catches none of this.

Specification gaming on verified math

The gaps in a verifier are exactly what gets exploited. We keep a growing set of pairs that satisfy the checker and two frontier LLM judges yet fail human review: 59 so far, sorted into named failure classes. Three of those classes now have mechanical detectors that run on everything we score.

The instrument itself

An evaluator you can’t calibrate is an opinion. Ours is measured against 886 frozen human verdicts: flag precision between 76.7 and 100 percent per corpus, recall 78.3 percent against expert corrections, and the false-positive classes named and published. Counterexample-backed flags have been upheld by human review 139 times out of 139. The instrument in full, pricing included, is on the eval suite page.

The strength-structured lattice

Where this goes next. Proofs organized in a lattice ordered by logical strength, not pass/fail, so a model’s verified reasoning can be measured for how much it proved. It also exposes a failure mode we expect to matter: models that train on self-verified math drifting toward weaker statements that are easier to prove.