Claude Opus 4.6 (Thinking) vs Spark 1.2 vs Claude Opus 5 (Thinking)
Side-by-side benchmark scores, pricing, and specifications
Specifications
| Specification | |||
|---|---|---|---|
| Provider | Anthropic | Meta | Anthropic |
| Variant | 4.6 Thinking | Spark 1.2 | Thinking |
| Input price | $5.00/1M | — | $5.00/1M |
| Output price | $25.00/1M | — | $25.00/1M |
| Context window | 1.0M | — | 1.0M |
| Benchmark | Comparison | Claude Opus 4.6 (Thinking) | Spark 1.2 | Claude Opus 5 (Thinking) |
|---|---|---|---|---|
CompositeQuality Score | 97.2%#15 | 104.4%#6 | 107.8%#4 | |
Human preferenceArena ELO | 1,505#1 | 1,500#3 | 1,493#6 | |
General reasoningLiveBench | 74.5%#32 | 78.0%#9 | 80.1%#5 | |
Scientific knowledgeGPQA Diamond | 91.3%#11 | — | — | |
Academic reasoningHLE | 40.0%#9 | — | — | |
Commonsense reasoningSimpleBench | 67.6%#11 | 74.5%#7 | 80.6%#2 | |
MathematicsAIME 2025 | 95.6%#2 | — | — | |
Graduate scienceGSO | 41.2%#3 | — | — | |
Agentic codingSWE-Bench Verified | 80.8%#5 | — | — | |
Agentic tool useTau-Bench | 91.9%#1 | — | — | |
Agentic terminalTerminal-Bench | 65.4%#5 | — | — | |
Visual reasoningARC-AGI-2 | 68.8%#12 | — | 88.3%#3 |
Scores represent the best available variant for each model. Higher is better unless otherwise noted. Bars show relative performance within each benchmark.