Written by eval/export_results.py from the raw runs. The protocol is in
docs/results.md.
| File | Contents |
|---|---|
answers.json |
AIME 2025, AIME 2026 and HMMT February 2026: the scores, and every sample's answer, finish reason and length |
proofbench-basic.json |
IMO-ProofBench Basic: the blind grade of every single-call proof and of the swarm's proof |
proofbench-basic/<id>.md |
one page per problem: the problem, the swarm's proof and all 16 single-call proofs, with grades and reasons |
The model outputs, grades and scores here are released under the repository's Apache-2.0 license. The benchmark material they contain keeps its own license:
- The problem statements in
proofbench-basic/are from IMO-ProofBench (Luong et al., 2025), google-deepmind/superhuman, Copyright 2025 Google LLC, licensed under CC BY 4.0 and reproduced unchanged. - The
goldanswers inanswers.jsoncome from opencompass/AIME2025 (MIT) for AIME 2025, and from MathArena's aime_2026 and hmmt_feb_2026 datasets (Dekoninck et al., 2026), licensed under CC BY-NC-SA 4.0.
How to cite the benchmarks is in docs/results.md.