Ranking
Best AI models for reasoning
Reasoning leaderboard for GPT, Claude, Gemini, Grok, and open-weight models. Ranked with LLM Stats TrueSkill and GPQA.
FAQ
Frequently asked questions
What is the best AI model for GPQA reasoning right now?
Rank 1 on this reasoning board is the conservative TrueSkill lead from LLM Stats with GPQA as the named benchmark. Use it when you care about hard scientific reasoning, not when you need LiveCodeBench coding or WritingBench long-form quality.
What does GPQA measure on an LLM reasoning leaderboard?
GPQA is a Google-proof Q&A set of graduate-level biology, physics, and chemistry items. LLM Stats maps it onto this reasoning ranking. It is a proxy for expert science questions, not a test of tool use, long context, or creative writing.
How is conservative TrueSkill different from a raw GPQA percent?
A raw GPQA score is percent correct on that eval. Conservative TrueSkill is a Bayesian skill rating with a penalty for uncertainty. This page shows the conservative rating so a model with a thin sample does not outrank a stable lead. LLM Stats still holds the raw GPQA numbers on the source board.
How do Claude, GPT, and Gemini compare on GPQA reasoning?
Read the live table. Labs share the same conservative scale, with logos and organization names on each row. The order changes as LLM Stats refreshes, typically about hourly, so compare the current rank rather than a remembered chart from last week.
Where does the reasoning LLM leaderboard data come from?
The reasoning ranking is republished from the public LLM Stats GPQA board at llm-stats.com, including conservative TrueSkill where LLM Stats publishes it. zerouter does not run GPQA or the other evals behind the board. Scores typically refresh about hourly.