Ranking

Best AI models for tool calling

Tool-calling and agent ranking of AI models. Conservative TrueSkill and MCP Atlas scores from LLM Stats.

20 modelsMCP AtlasMuse Spark 1.1 · 88.1
Tool calling scores
Top published models on this board · Sep 8, 2026, 11:25 PM UTC.
Labs on this board
Share of the top twenty rows by organization.
  • Alibaba Cloud / Qwen Team3
  • Meta2
  • Moonshot AI2
  • ByteDance2
  • Other11
Score line
Conservative rating by lab color.
Tool calling ranking
Full table for this skill board. Rank 1 is the current conservative lead.
RankModelScore
1Muse Spark 1.188.1
2Kimi K384.2
3BYSeed 2.1 Pro83.8
4TEHy4 preview83.7
5Gemini 3.5 Flash83.6
6Claude Opus 4.882.2
7BYSeed 2.1 Turbo80.3
8TMInkling-Small79.6
9TEHy379.1
10Claude Opus 4.777.3
11GLM-5.276.8
12Qwen3.7 Max76.4
13TMInkling76.0
14Kimi K2.7 Code76.0
15Muse Glimmer-30B75.5
16GPT-5.575.3
17MiniMax M374.2
18Qwen3.6 Plus74.1
19DeepSeek-V4-Pro-Max73.6
20Qwen3.7-Plus73.2

FAQ

Frequently asked questions

What is the best AI model for tool calling and function calling?

Rank 1 on this board is the conservative TrueSkill lead from LLM Stats with MCP Atlas as the named agent and tool-use eval. Use it when you care about calling tools and MCP-style agents, not when you only need LiveCodeBench code generation.

What is MCP Atlas on an LLM agent ranking?

MCP Atlas is the tool-calling and Model Context Protocol eval LLM Stats attaches to this board. It is closer to function calling and agent tool use than to GPQA or AIME. zerouter republishes the scores and does not run MCP Atlas.

How is a tool-calling leaderboard different from a coding leaderboard?

Coding here is LiveCodeBench program synthesis. Tool calling is whether the model selects and uses tools correctly. A coding lead can still miss tool schemas, so agent builders should read this board instead of copying the coding rank.

Why do agent rankings change faster than chat popularity lists?

Tool-calling evals move when labs ship new function-calling stacks. This table follows LLM Stats, which typically refreshes about hourly. Chat arena votes are a different signal and are not what this tool-calling board reports.

Where does the tool-calling LLM leaderboard data come from?

The tool calling ranking is republished from the public LLM Stats MCP Atlas board at llm-stats.com, including conservative TrueSkill where LLM Stats publishes it. zerouter does not run MCP Atlas or the other evals behind the board. Scores typically refresh about hourly.