Ranking

Best AI models for research

Research ranking of AI models. Independent LLM Stats TrueSkill ratings and MMLU-Pro scores.

20 modelsMMLU-ProSakana Namazu · 90.3
Research scores
Top published models on this board · Sep 8, 2026, 11:25 PM UTC.
Labs on this board
Share of the top twenty rows by organization.
  • Alibaba Cloud / Qwen Team10
  • DeepSeek3
  • Sakana AI1
  • MiniMax1
  • Other5
Score line
Conservative rating by lab color.
Research ranking
Full table for this skill board. Rank 1 is the current conservative lead.
RankModelScore
1SASakana Namazu90.3
2Qwen3.7 Max89.6
3Qwen3.7-Plus88.5
4Qwen3.6 Plus88.5
5MiniMax M2.188.0
6Qwen3.5-397B-A17B87.8
7DeepSeek-V4-Pro-Max87.5
8Kimi K2.587.1
9BAERNIE 5.087.0
10Nemotron 3 Ultra (550B A55B)86.8
11Qwen3.5-122B-A10B86.7
12DeepSeek-V4-Flash-042386.4
13UPSolar Pro 486.3
14DeepSeek-V4-Flash-Max86.2
15Qwen3.6-27B86.2
16Qwen3.5-27B86.1
17Qwen3 Max Thinking85.7
18Qwen3.5-35B-A3B85.3
19Qwen3.6-35B-A3B85.2
20Gemma 4 31B85.2

FAQ

Frequently asked questions

What is the best AI model for research and MMLU-Pro?

Rank 1 on this research board is the conservative TrueSkill lead from LLM Stats with MMLU-Pro as the named benchmark. Use it for broad academic coverage, not as a substitute for GPQA graduate science items or WritingBench prose.

What does MMLU-Pro measure on an LLM research ranking?

MMLU-Pro is a harder, more robust follow-on to MMLU: multiple-choice questions across academic subjects with more options and less guessability. LLM Stats uses it here as the research skill signal. zerouter does not administer MMLU-Pro.

How does the research board differ from the reasoning board?

Research on this page is MMLU-Pro breadth. Reasoning is GPQA depth in biology, physics, and chemistry. Pick research for wide subject coverage and reasoning when the task looks like expert science questions.

Are open-weight models competitive on the research LLM leaderboard?

Yes, when LLM Stats publishes a research score for them. The table flags open versus closed weights so you can filter Qwen, DeepSeek, GLM, and similar labs against GPT, Claude, and Gemini on the same conservative scale.

Where does the research LLM leaderboard data come from?

The research ranking is republished from the public LLM Stats MMLU-Pro board at llm-stats.com, including conservative TrueSkill where LLM Stats publishes it. zerouter does not run MMLU-Pro or the other evals behind the board. Scores typically refresh about hourly.