SWE-bench & software engineering
Compare SWE-bench Verified, SWE-Pro, SWE-Atlas, and Terminal-Bench scores for coding-focused model selection.
The AskClash LLM and AI leaderboard ranks GPT-5.6 Sol, Grok 4.6, Claude Opus/Fable, Gemini 3.7 Flash, DeepSeek V4 Pro 0813, GLM-5.3, Kimi K3, and frontier LLMs by SWE-bench, GPQA, HLE, Terminal-Bench, AskClash ACB, token pricing, and coding-agent scores. Updated daily. Compare GPT vs Claude vs Gemini, Grok 4.6, DeepSeek V4 Pro 0813, open-weight models, and the newest frontier releases in one benchmark table.
These ranked rows are cached from the live AskClash leaderboard. Open any model for SWE-bench breakdowns, benchmark results, pricing, and comparison links.
| # | Model | Creator | Overall | SWE | Price in/out |
|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 90.4 | 96.0 | $5.00 / $25.0 |
| 2 | Claude Fable 5 | Anthropic | 85.1 | 95.5 | $10.0 / $50.0 |
| 3 | GLM-5.3 | Z.AI | 83.1 | — | $1.40 / $4.40 |
| 4 | Grok 4.6 | xAI | 79.1 | — | $2.00 / $6.00 |
| 5 | GPT-5.6 Sol | OpenAI | 79.0 | — | $5.00 / $30.0 |
| 6 | Kimi K3 | Moonshot AI | 77.4 | — | $3.00 / $15.0 |
| 7 | Qwen3.8 Max | Alibaba | 74.4 | — | $2.00 / $6.00 |
| 8 | Claude Opus 4.8 | Anthropic | 73.7 | 88.6 | $5.00 / $25.0 |
| 9 | GPT-5.5 | OpenAI | 71.6 | — | $5.00 / $30.0 |
| 10 | GPT-5.6 Terra | OpenAI | 71.5 | — | $2.50 / $15.0 |
| 11 | Gemini 3.7 Flash | 71.1 | — | $1.50 / $7.50 | |
| 12 | Claude Sonnet 5 | Anthropic | 69.3 | 85.2 | $3.00 / $15.0 |
| 13 | Grok 4.5 | xAI | 68.6 | — | $2.00 / $6.00 |
| 14 | Muse Spark 1.2 | Meta | 66.7 | — | $1.25 / $4.25 |
| 15 | Ornith-1.5-397B | Ornith AI | 66.3 | 86.0 | $0 / $0 |
| 16 | GPT-5.6 Luna | OpenAI | 64.1 | — | $1.00 / $6.00 |
| 17 | Muse Spark 1.1 | Meta | 63.7 | — | $1.25 / $4.25 |
| 18 | DeepSeek V4 Pro 0813 | DeepSeek | 60.6 | 80.6 | $0.43 / $0.87 |
| 19 | DeepSeek V4 Flash Vision Exp | DeepSeek | 58.2 | 79.0 | $0.14 / $0.28 |
| 20 | Qwen3.8-27B | Alibaba | 56.9 | — | $0.45 / $3.20 |
This page targets common LLM comparison queries: best coding LLM, SWE-bench leaderboard, Grok 4.6 vs GPT-5.6 Sol, DeepSeek V4 Pro 0813, GPT vs Claude vs Gemini, LLM API pricing comparison, and frontier model rankings.
Compare SWE-bench Verified, SWE-Pro, SWE-Atlas, and Terminal-Bench scores for coding-focused model selection.
Track GPQA, MATH-500, HLE, ARC-AGI-2, Tau2, and multimodal benchmarks like MMMU-Pro when providers publish them.
AskClash ACB and Real World Testing add hands-on quality signals on top of published benchmark tables so rankings reflect practical use, not only vendor cards.
Frontier proprietary models, adaptive variants, and major open-weight releases from OpenAI, Anthropic, Google, Meta, DeepSeek, Alibaba, Moonshot, Z.ai, and Cursor when benchmark data is available.
AskClash combines multiple public benchmark sources into one weighted table with pricing, context, access path, and AskClash RWT instead of showing a single vendor index alone.
Benchmark snapshots refresh from cached collector data. Reload the live leaderboard for the newest model rows and scores.
The interactive leaderboard loads in the browser. This crawlable page gives search engines stable copy for LLM leaderboard, benchmark comparison, and model ranking intent.