Language Models

Pick the right model for your workload. Composite quality scores, honest tradeoffs, and value rankings across 129 models from 17 providers, 71 of them open source.

The verdict

Claude Fable 5 (Thinking) is the highest-quality LLM we track right now at $40/1M weighted tokens.

Best value: DeepSeek V4 (V4 Flash Thinking) scores 92 vs 114 for $0.16/1M instead of $40/1M. Cheapest serious pick: Microsoft Phi-4 (Reasoning Plus) at $0.12/1M.

Pareto Frontier

Best picks before the next price jump.

Use this as the action layer: each row is the strongest quality pick under that weighted token budget.

Pricing calculator
  1. < $0.15
    Quality81n=2
    $0.14/1M
  2. < $0.50
    Quality92n=10
    $0.16/1M
  3. < $1
    Quality92n=11
    $0.76/1M
  4. < $2
    Quality96n=9
    $1.87/1M
  5. < $5
    Quality97n=2
    $2.19/1M
  6. < $10
    Quality101n=10
    $7.81/1M

Live data

Cost vs. quality, measured

Models on the dashed line are Pareto-optimal: no other model offers better performance for less money. Pick a benchmark to change the lens.

Quality With Uncertainty

The leaderboard now shows range, not just rank.

Dot is the current estimate. Thick bars show the 80% interval; pale whiskers show the wider 90% interval.

How we score
Model
90
98
105
113
120
  1. Tier 1
    Overlapping 80% ranges share a tier
  2. 2
  3. 6
  4. Anthropic logo
    Claude Fable 5 (Thinking)
    6
  5. Anthropic logo
    Claude Opus 5 (Thinking)
    Provisional
    4
  6. OpenAI logo
    GPT-5.6 Sol (Thinking)
    Provisional
    5
  7. 4
  8. 11
  9. 11
  10. 10
  11. Moonshot logo
    Kimi K3 (Thinking)
    Provisional
    4
  12. 14
  13. 15
  14. Google logo
    Gemini 3.5 Flash (3.5)
    7
  15. 11

Model family directory

Family rollups are sorted by value first. For exact variant-level API costs, use the pricing calculator.

ValueQuality adjusted for input + output cost. Higher = more performance per dollar.QualityComposite benchmark point estimate across 14 evals (GPQA, SWE-bench, MMLU…). Nearby scores can tie statistically.ELOHuman preference rating from LMArena. Useful, but style-biased.How we score →
Microsoft logo
Microsoft Phi-4Non-ThinkingOSS

Microsoft

Value
106
Quality
Capable76/ ELO 1256
Context
16K
Price /1M
$0.07in · $0.14out
MiniMax logo
M3ThinkingOSS

MiniMax

Value
101
Quality
Capable88/ ELO 1444
Context
1.0M
Price /1M
$0.30in · $1.20out
Anthropic logo
Claude Opus 5Thinking

Anthropic

Value
99
Quality
Frontier112
Context
Price /1M
$5in · $25out
Zhipu AI logo
GLM-5.2ThinkingOSS

Zhipu AI

Value
98
Quality
Strong92/ ELO 1469
Context
1.0M
Price /1M
$0.73in · $2.28out
Google logo
Gemini 3.5 FlashThinking

Google

Value
98
Quality
Strong99/ ELO 1476
Context
1.0M
Price /1M
$1.50in · $9out
MiniMax logo
M2.7Thinking

MiniMax

Value
96
Quality
Capable84/ ELO 1417
Context
205K
Price /1M
$0.26in · $1.02out
Zhipu AI logo
GLM-5.1ThinkingOSS

Zhipu AI

Value
95
Quality
Strong92/ ELO 1470
Context
205K
Price /1M
$0.97in · $3.04out
Zhipu AI logo
GLM-5ThinkingOSS

Zhipu AI

Value
94
Quality
Capable88/ ELO 1457
Context
205K
Price /1M
$0.95in · $2.55out
Google logo
Gemini 3.6 FlashThinking

Google

Value
93
Quality
Strong94/ ELO 1485
Context
Price /1M
$1.50in · $7.50out
Moonshot logo
Kimi K3Thinking

Moonshot

Value
92
Quality
Elite100/ ELO 1486
Context
1.0M
Price /1M
$3in · $15out
Anthropic logo
Claude Fable 5Thinking

Anthropic

Value
91
Quality
Frontier114/ ELO 1507
Context
Price /1M
$10in · $50out
Zhipu AI logo
GLM-4.7ThinkingOSS

Zhipu AI

Value
90
Quality
Capable81/ ELO 1442
Context
205K
Price /1M
$0.06in · $0.40out
OpenAI logo
GPT-5.6 SolThinking

OpenAI

Value
88
Quality
Elite107/ ELO 1485
Context
Price /1M
$5in · $30out
OpenAI logo
GPT-5.6 LunaThinking

OpenAI

Value
88
Quality
Capable87
Context
Price /1M
$1in · $6out
Mistral AI logo
Medium 3.5ThinkingOSS

Mistral AI

Value
87
Quality
Capable88/ ELO 1427
Context
262K
Price /1M
$1.50in · $7.50out
thinking-machines logo
InklingThinkingOSS

thinking-machines

Value
87
Quality
Capable84/ ELO 1445
Context
1.0M
Price /1M
$1in · $4.05out
Google logo
Gemini 3.5 Flash LiteThinking

Google

Value
87
Quality
Capable80/ ELO 1459
Context
Price /1M
$0.30in · $2.50out
OpenAI logo
GPT-5.6 TerraThinking

OpenAI

Value
86
Quality
Strong94
Context
Price /1M
$2.50in · $15out
Anthropic logo
Claude Sonnet 5Thinking

Anthropic

Value
86
Quality
Strong93/ ELO 1461
Context
Price /1M
$3in · $15out
xAI logo
Grok 3Mini

xAI

Value
81
Quality
Capable74
Context
131K
Price /1M
$0.30in · $0.50out
Qwen logo
QwQ 32BThinkingOSS

Qwen

Value
77
Quality
Capable71/ ELO 1336
Context
33K
Price /1M
$0.20in · $0.20out
MiniMax logo
M1Reasoning

MiniMax

Value
75
Quality
Capable70/ ELO 1364
Context
1.0M
Price /1M
$0.55in · $2.20out
Zhipu AI logo
GLM-4.5 AirThinking

Zhipu AI

Value
73
Quality
Capable68/ ELO 1373
Context
131K
Price /1M
$0.13in · $0.85out
OpenAI logo
o4Mini

OpenAI

Value
71
Quality
Capable78/ ELO 1390
Context
Price /1M
$4in · $16out
xAI logo
Grok 3Thinking

xAI

Value
69
Quality
Capable75/ ELO 1412
Context
131K
Price /1M
$3in · $15out
OpenAI logo
o3Mini

OpenAI

Value
68
Quality
Capable67/ ELO 1363
Context
200K
Price /1M
$1.10in · $4.40out
Google logo
Gemma 3n 4BNon-thinkingOSS

Google

Value
67
Quality
Capable50/ ELO 1318
Context
33K
Price /1M
$0.06in · $0.12out
OpenAI logo
GPT-4.1Non-thinking

OpenAI

Value
64
Quality
Capable65/ ELO 1414
Context
1.0M
Price /1M
$2in · $8out
Microsoft logo
Phi-4 MiniNon-ThinkingOSS

Microsoft

Value
64
Quality
Capable63
Context
128K
Price /1M
$0.08in · $0.35out
OpenAI logo
o1 MiniThinking

OpenAI

Value
56
Quality
Capable59/ ELO 1337
Context
128K
Price /1M
$1.10in · $4.40out
Cohere logo
Command ANon-thinking

Cohere

Value
50
Quality
Capable56/ ELO 1354
Context
256K
Price /1M
$2.50in · $10out
OpenAI logo
GPT-4oNon-thinking

OpenAI

Value
49
Quality
Capable56/ ELO 1443
Context
128K
Price /1M
$2.50in · $10out
OpenAI logo
GPT-4o MiniNon-thinking

OpenAI

Value
45
Quality
Capable50/ ELO 1318
Context
128K
Price /1M
$0.15in · $0.60out
Cohere logo
Command RNon-thinking

Cohere

Value
32
Quality
Capable41/ ELO 1250
Context
128K
Price /1M
$0.50in · $1.50out
Qwen logo
Qwen 3.7 PlusThinking

Qwen

Value
Quality
ELO 1460
Context
1.0M
Price /1M
$0.32in · $1.28out
OpenAI logo
GPT-4.5Non-thinking

OpenAI

Value
No pricing yet
Quality
ELO 1445
Context
Price /1M
No pricing yet
Anthropic logo
Mythos 5Thinking

Anthropic · +1 more variants

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Google logo
Gemini 2.0 ProNon-thinking

Google

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Google logo
Gemma 3 1B ITNon-thinkingOSS

Google

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Google logo
Gemma 4 12B ITThinkingOSS

Google

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Google logo
Gemma 4 E2B ITThinkingOSS

Google

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Google logo
Gemma 4 E4B ITThinkingOSS

Google

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

xAI logo
Grok Build0.1

xAI

Value
Quality
No score yet
Context
256K
Price /1M
$1in · $2out
O
InternVL3.5 8BThinkingOSS

opengvlab

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

T
Penguin-VL 8BThinkingOSS

tencent

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Microsoft logo
Phi-4 Mini FlashNon-ThinkingOSS

Microsoft

Value
No pricing yet
Quality
No score yet
Context
128K
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Microsoft logo
Phi-4 Multimodal InstructNon-ThinkingOSS

Microsoft

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Microsoft logo
Phi-4 Reasoning Vision 15BThinkingOSS

Microsoft

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Qwen logo
Qwen3 VL 2BThinkingOSS

Qwen

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Qwen logo
Qwen3 VL 4BThinkingOSS

Qwen

Value
No pricing yet
Quality
No score yet
Context
Price /1M
No pricing yet

Recently added; benchmarks and pricing not yet available.

Get the Newsletter

Practical notes on model choice, API costs, self-hosting, and when lighter systems are good enough. Occasional, not a daily digest.

Join the list