Correct isn’t the same as measurable.

Verifiable rewards were a breakthrough for math and code: a deterministic checker decides whether an answer is right, and no learned judge gets fooled. But the checker only enforces what it can see. A Lean proof can compile cleanly while proving the wrong theorem, and nothing in the toolchain will object. Measuring reasoning therefore needs a second axis the compiler ignores: faithfulness, whether the formal statement says what the mathematics says.

We didn’t argue this in the abstract. We measured it. The benchmarks the field trains against circulated for years carrying defects nobody had filed. Frontier LLM judges, the standard fallback, passed 59 formalizations that human review rejected. Even expert correction passes introduced ten new errors while fixing others. Every layer of the current stack leaks, which is why our screen ends with a person, and why we publish the instrument’s own error rates next to every number it produces. The measurements live on the research page.

  • Compiles ≠ faithful A Lean proof can typecheck while proving the wrong statement.
  • Judges leak too 59 pairs passed two frontier judges and failed human review.
  • Humans close the loop Every published flag is human-confirmed; 139 of 139 counterexample-backed flags upheld.