Ranking

AI model ranking

Independent LLM rankings for coding, math, reasoning, writing, research, tool calling, and long context. Sourced from LLM Stats TrueSkill and published benchmarks.

7 live boardsLLM StatsConservative scores across LiveCodeBench and six more skill boards.
Coding

GPT-6 Astra leads

Board

48.5

trueskill
Math

GPT-5.2 Pro leads

Board
Writing

Qwen3-235B-A22B-Thinking-2507 leads

Board

88.3

WritingBench

  1. Qwen3-235B-A22B…88.3
  2. Qwen3-Next-80B-…87.3
  3. Qwen3 VL 235B A…86.7
  4. Qwen3 VL 32B Th…86.2

Reasoning

GPT-6 Astra

96.0

  1. 1 GPT-6 Astra96.0
  2. 2 GPT-5.6 Sol94.6
  3. 3 Claude Mythos Preview94.6
  4. 4 Gemini 3.1 Pro94.3

Research

Sakana Namazu

90.3

  1. 1 Sakana Namazu90.3
  2. 2 Qwen3.7 Max89.6
  3. 3 Qwen3.7-Plus88.5
  4. 4 Qwen3.6 Plus88.5

Tool calling

Muse Spark 1.1

88.1

  1. 1 Muse Spark 1.188.1
  2. 2 Kimi K384.2
  3. 3 Seed 2.1 Pro83.8
  4. 4 Hy4 preview83.7

Long context

Qwen3.8 Max

66.3

  1. 1 Qwen3.8 Max66.3
  2. 2 Qwen3.5-397B-A17B63.2
  3. 3 Qwen3.6 Plus62.0
  4. 4 Nemotron 3 Ultra (550B A55B)61.9

LiveCodeBench · trueskill · 30 models

RankModelScore
1GPT-6 Astra48.5
2GPT-5.6 Sol46.0
3Claude Fable 545.8
4Claude Mythos Preview45.0
5Kimi K343.5
6GLM-5.343.1
7GPT-5.6 Terra42.4
8Claude Opus 542.0
9Muse Spark 1.341.9
10Claude Opus 4.841.9
11DeepSeek-V4-Pro-081341.4
12TEHy4 preview40.0
13Qwen3.8 Max39.9
14Gemini 3.8 Flash39.3
15GPT-5.538.7
16Grok 4.637.9
17Claude Sonnet 537.1
18Claude Opus 4.736.9
19GPT-5.6 Luna36.5
20Qwen3.7 Max36.3
21Qwen3.8-Flash-Next36.2
22Qwen3.8 Flash36.2
23Gemini 3.7 Flash35.8
24GLM-5.235.5
25DeepSeek-V4-Flash-Vision-Exp35.4
26Grok 4.535.4
27GLM-5.3-Flash34.9
28Muse Spark 1.134.9
29Kimi K2.633.8
30TEHy333.5

New models

Announced in the last 15 days.

16 releases

Input prices

USD per 1M input tokens. Wider bars cost more.

FAQ

Frequently asked questions

What is the best AI model ranking for coding, math, and writing in 2026?

There is no single winner across every job. zerouter splits the LLM leaderboard into seven skill boards so you can pick a model for coding, math, reasoning, writing, research, tool calling, or long context. Rank 1 on each board is the current conservative TrueSkill lead from LLM Stats, not a blended intelligence score.

Where does the zerouter LLM leaderboard get its scores from?

Rankings on this page are republished from LLM Stats (llm-stats.com). The hub covers coding on LiveCodeBench, math on AIME 2025, reasoning on GPQA, writing on WritingBench, research on MMLU-Pro, tool calling on MCP Atlas, and long context on LongBench v2. zerouter does not run the underlying evals. Scores typically refresh about hourly.

What is a conservative TrueSkill rating on an AI model leaderboard?

The number on each row is LLM Stats conservative TrueSkill, a lower-confidence bound rather than the raw mean. It is used so a model with a few noisy wins does not jump the board. Compare it as a skill rating on that category, not as a percent on one exam unless the board also shows a named benchmark such as AIME 2025 or GPQA.

How often is the AI model ranking updated?

zerouter refreshes the published LLM Stats boards about once an hour. New model announcements can land on the New models strip within hours. If a skill board is empty, LLM Stats had no live table for that category at the last pull.

Which benchmark is used for each AI skill board?

Coding uses LiveCodeBench, math uses AIME 2025, reasoning uses GPQA, writing uses WritingBench, research uses MMLU-Pro, tool calling uses MCP Atlas, and long context uses LongBench v2. The table header also shows the ranking method LLM Stats published, often TrueSkill.

How do I compare Claude, GPT, Gemini, and open-weight models on one leaderboard?

Open a skill board and read rank, organization, conservative rating, and the open or closed weights flag together. That is how this page compares Claude, GPT, Gemini, Grok, Kimi, Qwen, DeepSeek, and other labs without mixing chat popularity with a coding or math score.

Does zerouter run LiveCodeBench, AIME 2025, or GPQA evaluations?

No. zerouter republishes the public LLM Stats boards. LiveCodeBench, AIME 2025, GPQA, WritingBench, MMLU-Pro, MCP Atlas, and LongBench v2 are scored by their publishers and aggregated by LLM Stats. Use llm-stats.com for methodology detail.

How are LLM input prices shown next to the model ranking?

Input prices are published USD per 1 million input tokens from the LLM Stats catalog when a rate exists. Switch the price board the same way you switch coding or writing. Missing prices mean LLM Stats did not publish an input rate, not that zerouter caller billing is free.