Language Models

Pick the right model for your workload. Composite quality scores, honest tradeoffs, and value rankings across 135 models from 17 providers, 75 of them open source.

The verdict

Claude Fable 5 (Thinking) is the highest-quality LLM we track right now at $40/1M weighted tokens.

Best value: DeepSeek V4 (V4 Pro Thinking) scores 95 vs 111 for $0.74/1M instead of $40/1M. Cheapest serious pick: DeepSeek V4 (V4 Flash Thinking) at $0.07/1M.

Pareto Frontier

Best picks before the next price jump.

Use this as the action layer: each row is the strongest quality pick under that weighted token budget.

Pricing calculator
  1. < $0.50
    Quality92n=11
    $0.07/1M
  2. < $2
    Quality95n=12
    $0.74/1M
  3. < $10
    Google logo
    Gemini 3.7 Flash (3.7)
    Provisional
    Quality105n=3
    $3.00/1M
  4. < $25
    Anthropic logo
    Claude Opus 5 (Thinking)
    Quality108n=6
    $20/1M

Live data

Cost vs. quality, measured

Models on the dashed line are Pareto-optimal: no other model offers better performance for less money. Pick a benchmark to change the lens.

Quality With Uncertainty

The leaderboard now shows range, not just rank.

Dot is the current estimate. Thick bars show the 80% interval; pale whiskers show the wider 90% interval.

How we score
Model
90
98
105
113
120
  1. Tier 1
    Overlapping 80% ranges share a tier
  2. 6
  3. 2
  4. Anthropic logo
    Claude Fable 5 (Thinking)
    6
  5. Anthropic logo
    Claude Opus 5 (Thinking)
    6
  6. Google logo
    Gemini 3.7 Flash (3.7)
    Provisional
    3
  7. 3
  8. OpenAI logo
    GPT-5.6 Sol (Thinking)
    Provisional
    5
  9. 4
  10. 11
  11. 11
  12. 2
  13. 10
  14. 14
  15. 11

Model family directory

Family rollups are sorted by value first. For exact variant-level API costs, use the pricing calculator.

ValueQuality adjusted for input + output cost. Higher = more performance per dollar.QualityComposite benchmark point estimate across 14 evals (GPQA, SWE-bench, MMLU…). Nearby scores can tie statistically.ELOHuman preference rating from LMArena. Useful, but style-biased.How we score →
Google logo
Gemini 3.7 FlashThinking

Google

Value
126
Quality
Elite105/ ELO 1490
Context
Price /1M
$0.75in · $3.75out
Anthropic logo
Claude Opus 5Thinking

Anthropic

Value
102
Quality
Elite108/ ELO 1493
Context
1.0M
Price /1M
$5in · $25out
Zhipu AI logo
GLM-5.3Thinking

Zhipu AI

Value
102
Quality
Strong95/ ELO 1483
Context
1.3M
Price /1M
$0.91in · $2.86out
Zhipu AI logo
GLM-5.3 FlashThinkingOSS

Zhipu AI

Value
101
Quality
Capable86/ ELO 1475
Context
1.3M
Price /1M
$0.09in · $0.30out
Zhipu AI logo
GLM-5.2ThinkingOSS

Zhipu AI

Value
99
Quality
Capable90/ ELO 1472
Context
1.0M
Price /1M
$0.55in · $1.74out
MiniMax logo
M3ThinkingOSS

MiniMax

Value
98
Quality
Capable86/ ELO 1441
Context
1.0M
Price /1M
$0.30in · $1.20out
OpenAI logo
GPT-5.6 LunaThinking

OpenAI

Value
98
Quality
Capable85/ ELO 1452
Context
1.1M
Price /1M
$0.20in · $1.20out
Zhipu AI logo
GLM-5.1ThinkingOSS

Zhipu AI

Value
96
Quality
Capable89/ ELO 1466
Context
205K
Price /1M
$0.97in · $3.04out
Zhipu AI logo
GLM-5ThinkingOSS

Zhipu AI

Value
95
Quality
Capable86/ ELO 1458
Context
205K
Price /1M
$0.60in · $1.92out
MiniMax logo
M2.7Thinking

MiniMax

Value
94
Quality
Capable82/ ELO 1415
Context
205K
Price /1M
$0.26in · $1.02out
OpenAI logo
GPT-5.6 SolThinking

OpenAI

Value
94
Quality
Elite104/ ELO 1483
Context
Price /1M
$4in · $20out
Qwen logo
Qwen 3.8 27BThinkingOSS

Qwen

Value
93
Quality
Capable84/ ELO 1437
Context
1.0M
Price /1M
$0.21in · $2.55out
Anthropic logo
Claude Fable 5Thinking

Anthropic

Value
92
Quality
Frontier110
Context
1.0M
Price /1M
$10in · $50out
Moonshot logo
Kimi K3Thinking

Moonshot

Value
92
Quality
Strong97/ ELO 1485
Context
1.0M
Price /1M
$1.70in · $8.50out
Google logo
Gemini 3.5 FlashThinking

Google

Value
91
Quality
Strong96/ ELO 1478
Context
1.0M
Price /1M
$1.50in · $9out
Google logo
Gemini 3.6 FlashThinking

Google

Value
91
Quality
Strong90/ ELO 1480
Context
Price /1M
$0.75in · $3.75out
Google logo
Gemini 3.5 Flash LiteThinking

Google

Value
87
Quality
Capable78/ ELO 1456
Context
Price /1M
$0.30in · $2.50out
Zhipu AI logo
GLM-4.7ThinkingOSS

Zhipu AI

Value
87
Quality
Capable78/ ELO 1442
Context
205K
Price /1M
$0.06in · $0.40out
Anthropic logo
Claude Sonnet 5Thinking

Anthropic

Value
86
Quality
Strong91/ ELO 1461
Context
1.0M
Price /1M
$2in · $10out
OpenAI logo
GPT-5.6 TerraThinking

OpenAI

Value
86
Quality
Strong91/ ELO 1466
Context
Price /1M
$2in · $12out
Microsoft logo
Microsoft Phi-4Non-ThinkingOSS

Microsoft

Value
85
Quality
Capable75/ ELO 1256
Context
16K
Price /1M
$0.07in · $0.14out
thinking-machines logo
InklingThinkingOSS

thinking-machines

Value
83
Quality
Capable82/ ELO 1440
Context
1.0M
Price /1M
$1in · $4.05out
Mistral AI logo
Medium 3.5ThinkingOSS

Mistral AI

Value
83
Quality
Capable86/ ELO 1426
Context
262K
Price /1M
$1.50in · $7.50out
xAI logo
Grok 3Mini

xAI

Value
81
Quality
Capable72
Context
131K
Price /1M
$0.30in · $0.50out
Qwen logo
QwQ 32BThinkingOSS

Qwen

Value
79
Quality
Capable71/ ELO 1336
Context
33K
Price /1M
$0.20in · $0.20out
MiniMax logo
M1Reasoning

MiniMax

Value
75
Quality
Capable70/ ELO 1364
Context
1.0M
Price /1M
$0.40in · $2.20out
Zhipu AI logo
GLM-4.5 AirThinking

Zhipu AI

Value
74
Quality
Capable67/ ELO 1373
Context
131K
Price /1M
$0.13in · $0.85out
OpenAI logo
o4Mini

OpenAI

Value
70
Quality
Capable76/ ELO 1391
Context
Price /1M
$4in · $16out
xAI logo
Grok 3Thinking

xAI

Value
68
Quality
Capable74/ ELO 1412
Context
131K
Price /1M
$3in · $15out
OpenAI logo
o3Mini

OpenAI

Value
67
Quality
Capable67/ ELO 1363
Context
200K
Price /1M
$1.10in · $4.40out
Microsoft logo
Phi-4 MiniNon-ThinkingOSS

Microsoft

Value
66
Quality
Capable62
Context
128K
Price /1M
$0.08in · $0.35out
OpenAI logo
GPT-4.1Non-thinking

OpenAI

Value
62
Quality
Capable64/ ELO 1414
Context
1.0M
Price /1M
$2in · $8out
OpenAI logo
o1 MiniThinking

OpenAI

Value
55
Quality
Capable59/ ELO 1337
Context
128K
Price /1M
$1.10in · $4.40out
Cohere logo
Command ANon-thinking

Cohere

Value
51
Quality
Capable56/ ELO 1354
Context
256K
Price /1M
$2.50in · $10out
Google logo
Gemma 3n 4BNon-thinkingOSS

Google

Value
49
Quality
Capable50/ ELO 1317
Context
33K
Price /1M
$0.06in · $0.12out
OpenAI logo
GPT-4o MiniNon-thinking

OpenAI

Value
48
Quality
Capable50/ ELO 1318
Context
128K
Price /1M
$0.07in · $0.30out
OpenAI logo
GPT-4oNon-thinking

OpenAI

Value
47
Quality
Capable54/ ELO 1443
Context
128K
Price /1M
$1.25in · $5out
Cohere logo
Command RNon-thinking

Cohere

Value
34
Quality
Capable41/ ELO 1250
Context
128K
Price /1M
$0.50in · $1.50out
Qwen logo
Qwen 3.7 PlusThinking

Qwen

Value
Quality
ELO 1456
Context
1.0M
Price /1M
$0.32in · $1.28out
OpenAI logo
GPT-4.5Non-thinking

OpenAI

Value
No pricing yet
Quality
ELO 1445
Context
Price /1M
No pricing yet
Anthropic logo
Mythos 5Thinking

Anthropic · +1 more variants

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Google logo
Gemini 2.0 ProNon-thinking

Google

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Google logo
Gemma 3 1B ITNon-thinkingOSS

Google

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Google logo
Gemma 4 12B ITThinkingOSS

Google

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Google logo
Gemma 4 E2B ITThinkingOSS

Google

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Google logo
Gemma 4 E4B ITThinkingOSS

Google

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

xAI logo
Grok Build0.1

xAI

Value
Quality
No score yet
Context
256K
Price /1M
$1in · $2out
O
InternVL3.5 8BThinkingOSS

opengvlab

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

T
Penguin-VL 8BThinkingOSS

tencent

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Microsoft logo
Phi-4 Mini FlashNon-ThinkingOSS

Microsoft

Value
No pricing yet
Quality
No score yet
Context
128K
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Microsoft logo
Phi-4 Multimodal InstructNon-ThinkingOSS

Microsoft

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Microsoft logo
Phi-4 Reasoning Vision 15BThinkingOSS

Microsoft

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Qwen logo
Qwen 3.8 2.4T A95BThinkingOSS

Qwen

Value
Quality
No score yet
Context
1.0M
Price /1M
$2in · $6out
Qwen logo
Qwen 3.8 Flash NextThinkingOSS

Qwen

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Qwen logo
Qwen3 VL 2BThinkingOSS

Qwen

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Qwen logo
Qwen3 VL 4BThinkingOSS

Qwen

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Get the Newsletter

Practical notes on model choice, API costs, self-hosting, and when lighter systems are good enough. Occasional, not a daily digest.

Join the list