Ranking
Best AI models for coding
Independent coding ranking of Claude, GPT, Gemini, Grok, Kimi, and open-weight models. Conservative TrueSkill scores from LLM Stats.
FAQ
Frequently asked questions
What is the best AI model for coding in 2026 on this leaderboard?
Rank 1 on this board is the current conservative TrueSkill lead for coding from LLM Stats, using LiveCodeBench as the named coding benchmark. Treat it as the best published coding model on this table today, not as a guarantee on your private repo or a SWE-Bench Verified substitute.
What does LiveCodeBench measure on an LLM coding ranking?
LiveCodeBench is a contamination-aware coding benchmark built from recent programming contest problems. LLM Stats uses it on this coding board so the ranking tracks live code generation rather than a static HumanEval snapshot. zerouter republishes that board and does not run LiveCodeBench.
How is this coding leaderboard different from SWE-Bench Verified?
SWE-Bench Verified scores software engineering patches against real GitHub issues. This page ranks models on LLM Stats coding TrueSkill with LiveCodeBench. A model that leads LiveCodeBench can still trail on SWE-Bench, so use this board for contest-style coding skill, not as a drop-in agent coding ranking.
Which open-weight models appear on the AI coding leaderboard?
Open-weight rows are marked Open in the weights column. Qwen, DeepSeek, GLM, Kimi, and other labs show up when LLM Stats publishes a coding score for them. Closed models such as GPT and Claude stay on the same table so you can compare API models against local weights.
Where does the coding LLM leaderboard data come from?
The coding ranking is republished from the public LLM Stats LiveCodeBench board at llm-stats.com, including conservative TrueSkill where LLM Stats publishes it. zerouter does not run LiveCodeBench or the other evals behind the board. Scores typically refresh about hourly.