Ranking
Best AI models for research
Research ranking of AI models. Independent LLM Stats TrueSkill ratings and MMLU-Pro scores.
FAQ
Frequently asked questions
What is the best AI model for research and MMLU-Pro?
Rank 1 on this research board is the conservative TrueSkill lead from LLM Stats with MMLU-Pro as the named benchmark. Use it for broad academic coverage, not as a substitute for GPQA graduate science items or WritingBench prose.
What does MMLU-Pro measure on an LLM research ranking?
MMLU-Pro is a harder, more robust follow-on to MMLU: multiple-choice questions across academic subjects with more options and less guessability. LLM Stats uses it here as the research skill signal. zerouter does not administer MMLU-Pro.
How does the research board differ from the reasoning board?
Research on this page is MMLU-Pro breadth. Reasoning is GPQA depth in biology, physics, and chemistry. Pick research for wide subject coverage and reasoning when the task looks like expert science questions.
Are open-weight models competitive on the research LLM leaderboard?
Yes, when LLM Stats publishes a research score for them. The table flags open versus closed weights so you can filter Qwen, DeepSeek, GLM, and similar labs against GPT, Claude, and Gemini on the same conservative scale.
Where does the research LLM leaderboard data come from?
The research ranking is republished from the public LLM Stats MMLU-Pro board at llm-stats.com, including conservative TrueSkill where LLM Stats publishes it. zerouter does not run MMLU-Pro or the other evals behind the board. Scores typically refresh about hourly.