LLM & AI Leaderboard

LLM and AI leaderboard for SWE-bench, coding agents, benchmarks, and API pricing.

The AskClash LLM and AI leaderboard ranks GPT-6 Astra, GPT-5.6 Sol, Grok 4.7, Grok 4.6, Claude Fable 5.1, Claude Opus 5, Gemini 3.8 Flash, Gemini 3.7 Flash, Qwen3.8 Max, DeepSeek V4.1 Flash, GLM-5.3-Flash, Qwen3.8 Flash Next, GLM-5.3, Kimi K3, and frontier LLMs by SWE-bench, GPQA, HLE, Terminal-Bench, AskClash ACB, token pricing, and coding-agent scores. Updated daily. Compare GPT vs Claude vs Gemini, Grok 4.6, DeepSeek V4 Pro 0813, open-weight models, and the newest frontier releases in one benchmark table.

#1 Claude Opus 5.5Current overall leader with score 80.4 across published benchmark signals and AskClash RWT.
SWE-bench & coding agentsTrack SWE-bench, SWE-Pro, SWE-Atlas, Terminal-Bench, and Artificial Analysis coding agent index scores side by side.
LLM pricing & contextSee input/output token pricing, context window, and a visit link for each model row.

Top LLM benchmark rankings

These ranked rows are cached from the live AskClash leaderboard. Open any model for SWE-bench breakdowns, benchmark results, pricing, and comparison links.

#ModelCreatorOverallSWEPrice in/out
1Claude Opus 5.5Anthropic80.4—$4.00 / $20.0
2Claude Fable 5.1Anthropic78.2—$10.0 / $50.0
3Claude Opus 5Anthropic77.196.0$5.00 / $25.0
4Claude Sonnet 5.5Anthropic76.8—$2.00 / $10.0
5Claude Fable 5Anthropic74.395.5$10.0 / $50.0
6GPT-6.1 SolOpenAI72.1—$2.00 / $10.0
7GPT-6 AstraOpenAI71.3—$10.0 / $50.0
8GPT-5.6 SolOpenAI69.6—$5.00 / $30.0
9Gemini 4 ArgonGoogle67.4—$2.00 / $10.0
10Qwen3.8 MaxAlibaba65.5—$2.00 / $6.00
11Muse Spark 1.3Meta63.6—$1.25 / $4.25
12GLM-5.3Z.AI63.4—$1.40 / $4.40
13GPT-5.6 TerraOpenAI61.4—$2.50 / $15.0
14Claude Opus 4.8Anthropic61.388.6$5.00 / $25.0
15SWE-2Cognition60.7—$0 / $0
16Claude Haiku 5.5Anthropic60.0—$0.10 / $0.50
17MiMo-V2.6-ProXiaomi59.8—$0.43 / $0.87
18Gemini 3.8 FlashGoogle59.7—$1.50 / $7.50
19GLM-5.3-FlashZ.AI59.6—$0.15 / $0.50
20DeepSeek V4.1 FlashDeepSeek59.1—$0.15 / $0.60

SWE-bench leaderboard

Top 10 models by published SWE-bench score. Claude Opus 5 leads at 96.0.

#ModelCreatorSWE-benchOverall rank
1Claude Opus 5Anthropic96.0#3
2Claude Fable 5Anthropic95.5#5
3Claude Opus 4.8Anthropic88.6#14
4Claude Sonnet 5Anthropic85.2#29
5Gemini 3.1 ProGoogle80.6#44
6MiniMax-M3MiniMax80.5#43
7Qwen3.7 MaxAlibaba80.4#34
8Claude Sonnet 4.6Anthropic80.2#54
9MiMo-V2.5-ProXiaomi78.9#51
10Qwen3.6 PlusAlibaba78.8#49

Best LLM for coding

Ranked by Coding Agent Index, the best LLM for coding right now is GPT-5.6 Sol (80.0), followed by Claude Fable 5 and GPT-5.6 Terra. For coding tools rather than models, see AI coding agents.

#ModelCreatorCoding Agent IndexOverall rank
1GPT-5.6 SolOpenAI80.0#8
2Claude Fable 5Anthropic77.2#5
3GPT-5.6 TerraOpenAI77.0#13
4Grok 4.5xAI76.4#31
5GPT-5.6 LunaOpenAI75.0#35
6GLM-5.2Z.AI74.4#30
7Claude Opus 4.8Anthropic72.5#14
8GPT-6 AstraOpenAI67.0#7
9Claude Opus 5Anthropic66.0#3
10GPT-6.1 SolOpenAI62.9#6

Benchmarks and search topics covered

This page targets common LLM comparison queries: best coding LLM, SWE-bench leaderboard, Grok 4.6 vs GPT-5.6 Sol, DeepSeek V4 Pro 0813, GPT vs Claude vs Gemini, LLM API pricing comparison, and frontier model rankings.

SWE-bench & software engineering

Compare SWE-bench Verified, SWE-Pro, SWE-Atlas, and Terminal-Bench scores for coding-focused model selection.

Reasoning & knowledge

Track GPQA, MATH-500, HLE, ARC-AGI-2, Tau2, and multimodal benchmarks like MMMU-Pro when providers publish them.

AskClash ACB & RWT

AskClash ACB and Real World Testing add hands-on quality signals on top of published benchmark tables so rankings reflect practical use, not only vendor cards.

Frequently asked questions

Which models are on the leaderboard?

Frontier proprietary models, adaptive variants, and major open-weight releases from OpenAI, Anthropic, Google, Meta, DeepSeek, Alibaba, Moonshot, Z.ai, and Cursor when benchmark data is available.

Is this the same as Artificial Analysis or LMSYS?

AskClash combines multiple public benchmark sources into one weighted table with pricing, context, access path, and AskClash RWT instead of showing a single vendor index alone.

How often do rankings update?

Benchmark snapshots refresh from cached collector data. Reload the live leaderboard for the newest model rows and scores.

The interactive leaderboard loads in the browser. This crawlable page gives search engines stable copy for LLM leaderboard, benchmark comparison, and model ranking intent.