Lingkai Kong, Zijian Wu, Yuzhe Gu, Haiteng Zhao, Zhouqi Hua, Wenyong Huang,
Shuang Sun, Zhicheng Xiong, Xiaotian Zhang, Shuya Zhao, Yan Wang, Disheng Xu,
Wenwei Zhang, Kai Chen
AdvancedMathBench evaluates whether language models can construct and verify natural-language proofs in advanced mathematics. It covers undergraduate-level (UG) and qualifying-examination-level (QE) mathematics, with problems drawn from examinations, mathematics competitions, and textbooks.
- ProverBench contains 245 expert-reviewed problems for proof generation.
- VerifierBench contains 888 proof trajectories with full-chain expert annotations, including fatal errors, recoverable errors, and explanations.
- AutoVerifier is a trained proof verifier for automated ProverBench evaluation. The default pessimistic protocol accepts a proof only when all eight judgments accept it. VerifierBench additionally uses meta-verification to assess whether a model's error analysis agrees with expert annotations.
Figure 1. Model performance on AdvancedMathBench. Left: proof-generation performance compared with HMMT and USAMO. Right: meta-verification TPR and TNR; marker labels show Balanced F1.
Figure 2. Overview of benchmark construction and the automatic verification pipeline.
| Benchmark | Evaluation unit | UG (ugd) |
QE (qe) |
Total |
|---|---|---|---|---|
| ProverBench | Problem | 200 | 45 | 245 |
| VerifierBench | Annotated proof | 168 | 720 | 888 |
The overview and results below follow the final manuscript dated September 28, 2026. The initial arXiv version describes an earlier ProverBench revision; see release notes for version details.
Main results on ProverBench and VerifierBench from the final manuscript, in percent. ProverBench uses pessimistic proof acceptance. VerifierBench reports both correct/incorrect polarity (Rough) and agreement with expert error analyses (Meta-Verification). Bal. F1 is the harmonic mean of TPR and TNR.
| Model | ProverBench | VerifierBench | ||||||
|---|---|---|---|---|---|---|---|---|
| UG | QE | Rough | Meta-Verification | |||||
| TPR | TNR | Bal. F1 | TPR | TNR | Bal. F1 | |||
| Proprietary Models | ||||||||
| GPT-5.5-xhigh | 64.5 | 48.9 | 78.9 | 73.3 | 76.0 | 78.9 | 55.1 | 64.9 |
| GPT-5.5-high | 53.3 | 46.1 | 78.4 | 73.8 | 76.0 | 78.3 | 53.6 | 63.6 |
| GPT-5.2 | 53.0 | 26.7 | 66.9 | 65.0 | 65.9 | 66.9 | 50.8 | 57.7 |
| Gemini-3.1-Pro-Preview | 46.5 | 17.8 | 94.0 | 49.0 | 64.4 | 94.0 | 39.1 | 55.2 |
| Claude-Opus-4.8 | 59.0 | 40.0 | 96.4 | 37.9 | 54.4 | 93.8 | 35.0 | 51.0 |
| Open-source Models | ||||||||
| DeepSeek-V4-Pro | 54.0 | 40.0 | 78.1 | 70.6 | 74.1 | 78.1 | 55.8 | 65.1 |
| Qwen3.5-397B-A17B | 40.0 | 33.5 | 91.5 | 55.2 | 68.9 | 91.5 | 43.0 | 58.5 |
| Kimi-K2.6 | 48.0 | 20.0 | 80.8 | 63.1 | 70.9 | 80.8 | 50.3 | 62.0 |
| GLM-5.2 | 44.5 | 28.9 | 82.5 | 66.9 | 73.9 | 82.3 | 51.5 | 63.3 |
| gpt-oss-120b | 20.5 | 2.2 | 95.2 | 38.7 | 55.0 | 95.3 | 32.0 | 47.9 |
| Intern-S2-Preview-35B | 27.0 | 16.7 | 95.2 | 38.7 | 55.0 | 95.2 | 30.7 | 46.4 |
Strong proof generation does not necessarily imply strong proof verification. Binary verdict accuracy can also hide incorrect error explanations. See the results notes for metric definitions and reproduction scope.
Python 3.10+ is required. Run the following commands from this repository's root.
The API evaluator has no internal-repository dependencies and uses the Python
standard library; the hub extra installs the Hugging Face download tools.
python -m pip install -e '.[hub]'
bash scripts/download_data.shThis downloads both benchmark files at a pinned revision and checks their
SHA256 hashes. If access requires authentication, run hf auth login first
with an authorized account. See data documentation.
Configure an API prover and an AutoVerifier endpoint:
cp configs/prover_api.example.json configs/prover_api.local.json
export POLICY_API_KEY="YOUR_API_KEY"
# Edit the model names and endpoint URLs in the local config.
python -m advancedmathbench prover \
--data data/proverbench/test.jsonl \
--config configs/prover_api.local.json \
--output outputs/prover \
--samples 4 --verifier-repeats 8 --dry-run--dry-run validates inputs and shows the call budget without calling models.
Remove it to run evaluation. Each of the four generated proofs per problem
receives eight independent verifier judgments. The example requires a deployed
AutoVerifier; it does not start a model server.
Configure the model to evaluate as a verifier:
cp configs/verifier_api.example.json configs/verifier_api.local.json
export VERIFIER_API_KEY="YOUR_API_KEY"
# Edit the model name and endpoint URL in the local config.
python -m advancedmathbench verifier \
--data data/verifierbench/test.jsonl \
--config configs/verifier_api.local.json \
--output outputs/verifier --verifier-repeats 1 --dry-runRemove --dry-run to run. For explanation-quality evaluation, use
configs/verifier_meta.example.json with --meta; the manuscript uses
gpt-oss-120b as the meta-verifier.
See the evaluation guide for meta-verification, local model loading, domain filters, outputs, and resuming. The metrics reference defines scoring and invalid-output handling. CLI scores are on 0–1; paper tables use percentages. Example settings are not a turnkey reproduction of paper results.
The model repository contains the checkpoint, tokenizer, and proof-verification prompt. To download the pinned checkpoint separately:
bash scripts/download_verifier.shThe checkpoint is approximately 68 GiB. Optional local inference uses public Transformers; see the setup and compatibility notes before loading it. Full GPU inference has not been validated by the offline tests.
advancedmathbench/ Standalone evaluator and bundled task prompts
assets/ Paper figures in SVG format
configs/ API and local-model configuration examples
data/ Dataset download instructions and checksums
docs/ Evaluation, metrics, results, and release notes
scripts/ Pinned dataset and checkpoint download helpers
tests/ Offline tests; no credentials or model weights needed
CITATION.cff Machine-readable paper citation
citation.bib BibTeX citation
Run tests with python -m unittest discover -s tests -v.
See validation for the tested scope.
No code license is declared at this time. Dataset and model terms are separate; see licensing notes and the respective Hugging Face repositories.
If you use AdvancedMathBench, please cite the paper. The citation below follows the current arXiv author list. Also available as BibTeX and CITATION.cff.
@misc{kong2026advancedmathbenchbenchmarksuiteadvanced,
title = {AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification},
author = {Lingkai Kong and Zijian Wu and Yuzhe Gu and Haiteng Zhao and Wenyong Huang and Shuang Sun and Zhicheng Xiong and Xiaotian Zhang and Shuya Zhao and Yan Wang and Disheng Xu and Wenwei Zhang and Kai Chen},
year = {2026},
eprint = {2607.11849},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
doi = {10.48550/arXiv.2607.11849},
url = {https://arxiv.org/abs/2607.11849}
}