Mathematics is being automated.
We check the math.
Millennium Research builds human-verified evaluations and certified datasets for formal mathematics. When a model translates a theorem into Lean, we measure whether the formal statement says what the mathematics says. The compiler can’t tell you. It turns out frontier LLM judges often can’t either.
The field measures AI on benchmarks whose statements don’t always say what the math says. We ran the census that proves it, by hand, one flag at a time.
Every published flag survives a person
Across every audit to date, 156 of 170 sampled flags survived human review. Per-corpus precision is published with its confidence interval next to each result.
Counterexample flags · cumulative record 139 / 139The evidence class that has never missed
When the screen produces concrete values where the hypotheses hold and the conclusion fails, human review has upheld it every time to date: three benchmarks and four autoformalizers.
Recall · against expert corrections 78.3%Most of what experts fix, the screen already flagged
Measured against the independent miniF2F v2 correction effort: of 92 confirmed defect corrections, 72 were statements our screen had flagged on its own.
Public record · benchmarks and autoformalizers 76 defectsConfirmed, documented, and filed
Defects in miniF2F, miniF2F v2, and ProofNet#, each with per-item evidence, plus a four-system audit of open autoformalizers. All three filings are public.
The faithfulness eval
Two frontier LLM judges under strict consensus, a numeric counterexample probe, and a detection ladder that shows, pair by pair, what the compiler missed, what single judges missed, and what only a human caught.
Calibrated, not claimed. Judge error rates are measured against 886 frozen human verdicts and published with confidence intervals.
- independent judges per pair2, strict consensus
- evidence per flagback-translations + counterexamples
- calibration set886 frozen human verdicts
The certified dataset
Informal–formal statement pairs from open-licensed mathematics, machine-verified in Lean 4 and certified faithful by human reviewers. Automated screens can only reject. Only a human can certify.
Firewalled by construction. The certified product never ships proof text. License-gated at ingest. Audited before every publish, with a tamper-evident history.
- verification oracleLean 4 + mathlib
- certificationhuman-only
- licenses in corpusCC0 / CC-BY
The endgame is a prover that moves open problems.
Evaluations and certified data are not the destination. They are the supply chain for machines that do mathematics and can prove it. Faithful formalization is the prerequisite; verified search is the method; a bound on an open problem, machine-checked end to end, is the kind of result we intend to put our name on.