Ranking

Best AI models for long context

Long-context ranking of AI models. LLM Stats TrueSkill and LongBench v2 scores for models that hold large windows.

18 modelsLongBench v2Qwen3.8 Max · 66.3
  1. #1

    Qwen3.8 Max

    Alibaba Cloud / Qwen Team

    66.366.3
  2. #2

    Qwen3.5-397B-A17B

    Alibaba Cloud / Qwen Team

    63.263.2
  3. #3

    Qwen3.6 Plus

    Alibaba Cloud / Qwen Team

    62.062.0
Score line
Conservative rating by lab color. · Sep 8, 2026, 11:25 PM UTC
Labs on this board
Share of the top twenty rows by organization.
  • Alibaba Cloud / Qwen Team11
  • MiniMax2
  • NVIDIA1
  • Microsoft1
  • Other3
Long context ranking
Full table for this skill board. Rank 1 is the current conservative lead.
RankModelScore
1Qwen3.8 Max66.3
2Qwen3.5-397B-A17B63.2
3Qwen3.6 Plus62.0
4Nemotron 3 Ultra (550B A55B)61.9
5MiniMax M1 80K61.5
6MAI-Thinking-161.0
7Kimi K2.561.0
8MiniMax M1 40K61.0
9Qwen3.5-27B60.6
10Qwen3 Max Thinking60.6
11XIMiMo-V2-Flash60.6
12Qwen3.5-122B-A10B60.2
13Qwen3.5-35B-A3B59.0
14Qwen3.5-9B55.2
15Qwen3.5-4B50.0
16DeepSeek-V348.7
17Qwen3.5-2B38.7
18Qwen3.5-0.8B26.1

FAQ

Frequently asked questions

What is the best long context AI model in 2026 on this board?

Rank 1 is the conservative TrueSkill lead from LLM Stats with LongBench v2 as the named long-context eval. A large advertised context window is not the same as a high LongBench v2 score, so read this table before you pick a 1M-token model on window size alone.

What does LongBench v2 measure on a long-context LLM ranking?

LongBench v2 tests whether a model can use information spread through a long document, not merely accept a long prompt. LLM Stats uses it for this skill board. zerouter does not run LongBench v2.

How is context window size different from long-context ranking quality?

Context window is the token limit the API accepts. Long-context quality is whether the model still retrieves and reasons over that span. This ranking is the quality score. Check the provider's window separately if you need a hard token ceiling.

How should I read the long-context AI leaderboard?

Start with the lead strip, then the score line colored by lab, then the full table. Conservative ratings on LongBench v2 are the sort key. Input prices, when published, are USD per 1M input tokens and are not the long-context score.

Where does the long-context LLM leaderboard data come from?

The long context ranking is republished from the public LLM Stats LongBench v2 board at llm-stats.com, including conservative TrueSkill where LLM Stats publishes it. zerouter does not run LongBench v2 or the other evals behind the board. Scores typically refresh about hourly.