Ranking

Best AI models for math

Independent math ranking of frontier LLMs. Conservative TrueSkill and AIME scores from LLM Stats.

20 modelsAIME 2025GPT-5.2 Pro · 100.0
Math scores
Top published models on this board · Sep 8, 2026, 11:25 PM UTC.
Labs on this board
Share of the top twenty rows by organization.
  • OpenAI6
  • Google2
  • Moonshot AI2
  • Sarvam AI2
  • Other8
Math ranking
Full table for this skill board. Rank 1 is the current conservative lead.
RankModelScore
1GPT-5.2 Pro100.0
2GPT-5.2100.0
3Gemini 3 Pro100.0
4Kimi K2-Thinking-0905100.0
5Grok-4 Heavy100.0
6Claude Opus 4.699.8
7Gemini 3 Flash99.7
8MELongCat-Flash-Thinking-260199.6
9GPT-5.1 High99.6
10Nemotron 3 Nano (30B A3B)99.2
11GPT OSS 20B High98.7
12GPT-5.1 Medium98.4
13BYSeed 2.0 Pro98.3
14STStep-3.5-Flash97.3
15MAI-Thinking-197.0
16SASarvam-105B96.7
17SASarvam-30B96.7
18GPT-5.1 Codex High96.7
19Kimi K2.596.1
20DeepSeek-V3.2-Speciale96.0

FAQ

Frequently asked questions

What is the best AI model for AIME 2025 math problems?

The math board ranks models with LLM Stats conservative TrueSkill and AIME 2025 as the named contest. A 100.0 lead means that model is at the top of the published AIME 2025 table LLM Stats is showing, not that it will solve every unpublished contest problem.

How is the math LLM ranking scored?

Scores are conservative TrueSkill ratings tied to AIME 2025. AIME is a short-answer high school contest used as a hard math probe for frontier models. Read the rating as contest math skill on this board, separate from GPQA science reasoning or LiveCodeBench coding.

Is a 100.0 score a perfect AIME 2025 result?

On this page 100.0 is the scaled conservative rating LLM Stats published for the current math lead. Several models can sit at 100.0 when the board is capped. Check rank order and the source AIME 2025 table on LLM Stats if you need the raw contest breakdown.

How does the math board differ from the GPQA reasoning ranking?

AIME 2025 is contest mathematics. GPQA is graduate-level science questions used on the reasoning board. A model can lead math and trail reasoning, so pick the board that matches the work: symbolic contests here, scientific QA on /ranking/reasoning.

Where does the math LLM leaderboard data come from?

The math ranking is republished from the public LLM Stats AIME 2025 board at llm-stats.com, including conservative TrueSkill where LLM Stats publishes it. zerouter does not run AIME 2025 or the other evals behind the board. Scores typically refresh about hourly.