Ranking
Best AI models for math
Independent math ranking of frontier LLMs. Conservative TrueSkill and AIME scores from LLM Stats.
FAQ
Frequently asked questions
What is the best AI model for AIME 2025 math problems?
The math board ranks models with LLM Stats conservative TrueSkill and AIME 2025 as the named contest. A 100.0 lead means that model is at the top of the published AIME 2025 table LLM Stats is showing, not that it will solve every unpublished contest problem.
How is the math LLM ranking scored?
Scores are conservative TrueSkill ratings tied to AIME 2025. AIME is a short-answer high school contest used as a hard math probe for frontier models. Read the rating as contest math skill on this board, separate from GPQA science reasoning or LiveCodeBench coding.
Is a 100.0 score a perfect AIME 2025 result?
On this page 100.0 is the scaled conservative rating LLM Stats published for the current math lead. Several models can sit at 100.0 when the board is capped. Check rank order and the source AIME 2025 table on LLM Stats if you need the raw contest breakdown.
How does the math board differ from the GPQA reasoning ranking?
AIME 2025 is contest mathematics. GPQA is graduate-level science questions used on the reasoning board. A model can lead math and trail reasoning, so pick the board that matches the work: symbolic contests here, scientific QA on /ranking/reasoning.
Where does the math LLM leaderboard data come from?
The math ranking is republished from the public LLM Stats AIME 2025 board at llm-stats.com, including conservative TrueSkill where LLM Stats publishes it. zerouter does not run AIME 2025 or the other evals behind the board. Scores typically refresh about hourly.