Ranking
AI model ranking
Independent LLM rankings for coding, math, reasoning, writing, research, tool calling, and long context. Sourced from LLM Stats TrueSkill and published benchmarks.
Coding
GPT-6 Astra
48.5
Math
GPT-5.2 Pro
100.0
Reasoning
GPT-6 Astra
96.0
Writing
Qwen3-235B-A22B-Thinking-2507
88.3
Research
SASakana Namazu
90.3
Tool calling
Muse Spark 1.1
88.1
Long context
Qwen3.8 Max
66.3
Reasoning
GPT-6 Astra
96.0
- 1 GPT-6 Astra96.0
- 2 GPT-5.6 Sol94.6
- 3 Claude Mythos Preview94.6
- 4 Gemini 3.1 Pro94.3
Research
Sakana Namazu
90.3
- 1 Sakana Namazu90.3
- 2 Qwen3.7 Max89.6
- 3 Qwen3.7-Plus88.5
- 4 Qwen3.6 Plus88.5
Tool calling
Muse Spark 1.1
88.1
- 1 Muse Spark 1.188.1
- 2 Kimi K384.2
- 3 Seed 2.1 Pro83.8
- 4 Hy4 preview83.7
Long context
Qwen3.8 Max
66.3
- 1 Qwen3.8 Max66.3
- 2 Qwen3.5-397B-A17B63.2
- 3 Qwen3.6 Plus62.0
- 4 Nemotron 3 Ultra (550B A55B)61.9
LiveCodeBench · trueskill · 30 models
| Rank | Model | Score |
|---|---|---|
| 1 | 48.5 | |
| 2 | 46.0 | |
| 3 | 45.8 | |
| 4 | 45.0 | |
| 5 | 43.5 | |
| 6 | GLM-5.3 | 43.1 |
| 7 | 42.4 | |
| 8 | 42.0 | |
| 9 | Muse Spark 1.3 | 41.9 |
| 10 | 41.9 | |
| 11 | DeepSeek-V4-Pro-0813 | 41.4 |
| 12 | TEHy4 preview | 40.0 |
| 13 | Qwen3.8 Max | 39.9 |
| 14 | 39.3 | |
| 15 | 38.7 | |
| 16 | 37.9 | |
| 17 | 37.1 | |
| 18 | 36.9 | |
| 19 | 36.5 | |
| 20 | Qwen3.7 Max | 36.3 |
| 21 | Qwen3.8-Flash-Next | 36.2 |
| 22 | Qwen3.8 Flash | 36.2 |
| 23 | 35.8 | |
| 24 | GLM-5.2 | 35.5 |
| 25 | DeepSeek-V4-Flash-Vision-Exp | 35.4 |
| 26 | 35.4 | |
| 27 | GLM-5.3-Flash | 34.9 |
| 28 | Muse Spark 1.1 | 34.9 |
| 29 | 33.8 | |
| 30 | TEHy3 | 33.5 |
New models
Announced in the last 15 days.
GPT-6 Astra
OpenAI
Sep 3$1.00- IN
Ling 3.0 Flash Fin
InclusionAI
Sep 3$0.01 Gemini 3.8 Flash
Google
Sep 2$0.75Gemini 3.8 Flash Cyber
Google
Sep 2—Muse Spark 1.3
Meta
Sep 2$0.002
Claude Fable 5.1
Anthropic
Sep 1$0.25
Claude Mythos 5.1
Anthropic
Sep 1—- TE
Hy4 preview
Tencent
Aug 28— Gemini Omni 1.1 Flash
Google
Aug 27$1.50Parse
Cohere
Aug 27—- CA
Sonic 3.6
Cartesia
Aug 27$10 Gemini 3.5 Transcribe
Google
Aug 26$2.00Gemini 3.5 Transcribe Live
Google
Aug 26$3.50GLM-5.3-Flash
Zhipu AI
Aug 26$0.03Qwen3.8 Flash
Alibaba Cloud / Qwen Team
Aug 26$0.02Qwen3.8-Flash-Next
Alibaba Cloud / Qwen Team
Aug 26—
Input prices
USD per 1M input tokens. Wider bars cost more.
FAQ
Frequently asked questions
What is the best AI model ranking for coding, math, and writing in 2026?
There is no single winner across every job. zerouter splits the LLM leaderboard into seven skill boards so you can pick a model for coding, math, reasoning, writing, research, tool calling, or long context. Rank 1 on each board is the current conservative TrueSkill lead from LLM Stats, not a blended intelligence score.
Where does the zerouter LLM leaderboard get its scores from?
Rankings on this page are republished from LLM Stats (llm-stats.com). The hub covers coding on LiveCodeBench, math on AIME 2025, reasoning on GPQA, writing on WritingBench, research on MMLU-Pro, tool calling on MCP Atlas, and long context on LongBench v2. zerouter does not run the underlying evals. Scores typically refresh about hourly.
What is a conservative TrueSkill rating on an AI model leaderboard?
The number on each row is LLM Stats conservative TrueSkill, a lower-confidence bound rather than the raw mean. It is used so a model with a few noisy wins does not jump the board. Compare it as a skill rating on that category, not as a percent on one exam unless the board also shows a named benchmark such as AIME 2025 or GPQA.
How often is the AI model ranking updated?
zerouter refreshes the published LLM Stats boards about once an hour. New model announcements can land on the New models strip within hours. If a skill board is empty, LLM Stats had no live table for that category at the last pull.
Which benchmark is used for each AI skill board?
Coding uses LiveCodeBench, math uses AIME 2025, reasoning uses GPQA, writing uses WritingBench, research uses MMLU-Pro, tool calling uses MCP Atlas, and long context uses LongBench v2. The table header also shows the ranking method LLM Stats published, often TrueSkill.
How do I compare Claude, GPT, Gemini, and open-weight models on one leaderboard?
Open a skill board and read rank, organization, conservative rating, and the open or closed weights flag together. That is how this page compares Claude, GPT, Gemini, Grok, Kimi, Qwen, DeepSeek, and other labs without mixing chat popularity with a coding or math score.
Does zerouter run LiveCodeBench, AIME 2025, or GPQA evaluations?
No. zerouter republishes the public LLM Stats boards. LiveCodeBench, AIME 2025, GPQA, WritingBench, MMLU-Pro, MCP Atlas, and LongBench v2 are scored by their publishers and aggregated by LLM Stats. Use llm-stats.com for methodology detail.
How are LLM input prices shown next to the model ranking?
Input prices are published USD per 1 million input tokens from the LLM Stats catalog when a rate exists. Switch the price board the same way you switch coding or writing. Missing prices mean LLM Stats did not publish an input rate, not that zerouter caller billing is free.