Back to the overview

Unified ranking, lane-aware explanation

Why the ranking looks like this.

The public ranking is not a single vibe score. It orders measured entrants by a transparent overall formula while keeping Full / Agentic, SWE MVP, Hard Intelligence, cost, and reliability visible.

Ranked entrants27

27 with Hard Intelligence data

Current leaderDeepSeek V4.1 Flash

Overall 89.15

Score spread39.04

#1 to #27

FormulaLane mean

Full + SWE + published Hard Intelligence when measured

Data refreshSep 25, 2026

Static HTML plus public JSON

Figure 1 Overall ladder Every ranked entrant ordered by public overall score.
DeepSeek V4.1 Flash 89.15 rank #1
GPT‑5.6 Terra 86.72 rank #2
GPT‑5.5 85.05 rank #3
DeepSeek V4 Flash 84.35 rank #4
GPT‑5.6 Sol 83.91 rank #5
Claude Opus 4.8 82.74 rank #6
Claude Fable 5 82.08 rank #7
Gemini 3.5 Flash 82.02 rank #8
GLM 5.3 Flash 81.32 rank #9
GPT‑5.6 Luna 81.15 rank #10
Claude Opus 5.5 80.87 rank #11
Claude Fable 5.1 80.53 rank #12
GLM‑5.2 80.50 rank #13
Claude Sonnet 5 80.44 rank #14
DeepSeek V4 Pro 79.37 rank #15
GPT‑6 Astra 78.29 rank #16
Qwen3.7 Max 77.52 rank #17
MiniMax M3 77.19 rank #18
GPT‑6 Sol 76.09 rank #19
GPT‑6 Luna 74.99 rank #20
MiniMax M3 Direct Plus 70.51 rank #21
Kimi K2.7 Code 68.18 rank #22
Step 3.7 Flash 67.16 rank #23
NVIDIA Nemotron 3 Ultra 63.45 rank #24
Qwythos‑9B Claude Mythos Q8_0Local model 52.86 rank #26
Ornith‑1.0‑35B Q4_K_MLocal model 50.11 rank #27

Overall is a lane mean, not a hidden replacement for source measurements.

Tradeoff scatter maps

Each point is one tested model at the intersection of two public telemetry axes. Use the maps to read quality against cost, speed, and recorded token use. Runtime and token axes are normalized per item and show coverage in the hover cards; they are telemetry, not current overall score inputs.

Figure 2 Cost × overall Overall score against measured cost. Each point is one tested model; hover or focus a point for its card.
Names stay on the top 10; every visible point keeps its rank label and hover detail.
Cost × overall Each point is one tested model. Measured cost is plotted on the X-axis and Overall score is plotted on the Y-axis. $047.0 $5.23658.3 $10.47369.6 $15.70981.0 $20.94592.3 #1 - DS-V4.1f · DeepSeek V4.1 Flash · Measured cost $0.183 · Overall score 89.2 #1 - DS-V4.1f #2 - Terra · GPT‑5.6 Terra · Measured cost $1.867 · Overall score 86.7 #2 - Terra #3 - GPT5.5 · GPT‑5.5 · Measured cost $4.822 · Overall score 85.0 #3 - GPT5.5 #4 - DS-V4f · DeepSeek V4 Flash · Measured cost $0.290 · Overall score 84.4 #4 - DS-V4f #5 - Sol · GPT‑5.6 Sol · Measured cost $3.858 · Overall score 83.9 #5 - Sol #6 - Opus · Claude Opus 4.8 · Measured cost $6.115 · Overall score 82.7 #6 - Opus #7 - Fable · Claude Fable 5 · Measured cost $12.285 · Overall score 82.1 #7 - Fable #8 - Gemini · Gemini 3.5 Flash · Measured cost $1.646 · Overall score 82.0 #8 - Gemini #9 - GLM 5.3 Flash · GLM 5.3 Flash · Measured cost $0.117 · Overall score 81.3 #9 - GLM 5.3 Flash #10 - Luna · GPT‑5.6 Luna · Measured cost $1.000 · Overall score 81.1 #10 - Luna #11 - Opus 5.5 · Claude Opus 5.5 · Measured cost $7.829 · Overall score 80.9 #11 #12 - Claude Fable 5.1 · Claude Fable 5.1 · Measured cost $19.394 · Overall score 80.5 #12 #13 - GLM5.2 · GLM‑5.2 · Measured cost $2.768 · Overall score 80.5 #13 #14 - Sonnet · Claude Sonnet 5 · Measured cost $2.812 · Overall score 80.4 #14 #15 - DS-V4p · DeepSeek V4 Pro · Measured cost $0.335 · Overall score 79.4 #15 #16 - GPT‑6 Astra · GPT‑6 Astra · Measured cost $6.776 · Overall score 78.3 #16 #17 - Qwen3.7 · Qwen3.7 Max · Measured cost $0.906 · Overall score 77.5 #17 #18 - M3 · MiniMax M3 · Measured cost $0.182 · Overall score 77.2 #18 #19 - GPT‑6 Sol · GPT‑6 Sol · Measured cost $1.310 · Overall score 76.1 #19 #20 - GPT‑6 Luna · GPT‑6 Luna · Measured cost $0.070 · Overall score 75.0 #20 #21 - M3 Direct · MiniMax M3 Direct Plus · Measured cost $0.169 · Overall score 70.5 #21 #22 - Kimi · Kimi K2.7 Code · Measured cost $0.732 · Overall score 68.2 #22 #23 - Step · Step 3.7 Flash · Measured cost $0.494 · Overall score 67.2 #23 #24 - Nemotron · NVIDIA Nemotron 3 Ultra · Measured cost $0.489 · Overall score 63.4 #24 #25 - Gemma · Gemma4‑12B‑Coder Fable5/Composer2.5 Q4_K_M · Measured cost $0 · Overall score 57.6 #25 #26 - Qwythos · Qwythos‑9B Claude Mythos Q8_0 · Measured cost $0 · Overall score 52.9 #26 #27 - Ornith · Ornith‑1.0‑35B Q4_K_M · Measured cost $0 · Overall score 50.1 #27 Measured cost Overall score
Figure 3 Runtime × overall Overall score against seconds / timed item. Each point is one tested model; hover or focus a point for its card.
Names stay on the top 10; every visible point keeps its rank label and hover detail.
Runtime × overall Each point is one tested model. Seconds / timed item is plotted on the X-axis and Overall score is plotted on the Y-axis. 0.00s47.0 31.7s58.3 63.3s69.6 95.0s81.0 127s92.3 #1 - DS-V4.1f · DeepSeek V4.1 Flash · Seconds / timed item 15.1s · Overall score 89.2 #1 - DS-V4.1f #2 - Terra · GPT‑5.6 Terra · Seconds / timed item 5.75s · Overall score 86.7 #2 - Terra #3 - GPT5.5 · GPT‑5.5 · Seconds / timed item 19.6s · Overall score 85.0 #3 - GPT5.5 #4 - DS-V4f · DeepSeek V4 Flash · Seconds / timed item 22.9s · Overall score 84.4 #4 - DS-V4f #5 - Sol · GPT‑5.6 Sol · Seconds / timed item 13.3s · Overall score 83.9 #5 - Sol #6 - Opus · Claude Opus 4.8 · Seconds / timed item 10.4s · Overall score 82.7 #6 - Opus #7 - Fable · Claude Fable 5 · Seconds / timed item 12.4s · Overall score 82.1 #7 - Fable #8 - Gemini · Gemini 3.5 Flash · Seconds / timed item 9.13s · Overall score 82.0 #8 - Gemini #9 - GLM 5.3 Flash · GLM 5.3 Flash · Seconds / timed item 35.5s · Overall score 81.3 #9 - GLM 5.3 Flash #10 - Luna · GPT‑5.6 Luna · Seconds / timed item 8.02s · Overall score 81.1 #10 - Luna #11 - Opus 5.5 · Claude Opus 5.5 · Seconds / timed item 7.42s · Overall score 80.9 #11 #12 - Claude Fable 5.1 · Claude Fable 5.1 · Seconds / timed item 8.97s · Overall score 80.5 #12 #13 - GLM5.2 · GLM‑5.2 · Seconds / timed item 91.2s · Overall score 80.5 #13 #14 - Sonnet · Claude Sonnet 5 · Seconds / timed item 12.5s · Overall score 80.4 #14 #15 - DS-V4p · DeepSeek V4 Pro · Seconds / timed item 23.2s · Overall score 79.4 #15 #16 - GPT‑6 Astra · GPT‑6 Astra · Seconds / timed item 13.5s · Overall score 78.3 #16 #17 - Qwen3.7 · Qwen3.7 Max · Seconds / timed item 22.7s · Overall score 77.5 #17 #18 - M3 · MiniMax M3 · Seconds / timed item 27.5s · Overall score 77.2 #18 #19 - GPT‑6 Sol · GPT‑6 Sol · Seconds / timed item 12.1s · Overall score 76.1 #19 #20 - GPT‑6 Luna · GPT‑6 Luna · Seconds / timed item 13.4s · Overall score 75.0 #20 #21 - M3 Direct · MiniMax M3 Direct Plus · Seconds / timed item 21.9s · Overall score 70.5 #21 #22 - Kimi · Kimi K2.7 Code · Seconds / timed item 28.6s · Overall score 68.2 #22 #23 - Step · Step 3.7 Flash · Seconds / timed item 27.8s · Overall score 67.2 #23 #24 - Nemotron · NVIDIA Nemotron 3 Ultra · Seconds / timed item 70.7s · Overall score 63.4 #24 #25 - Gemma · Gemma4‑12B‑Coder Fable5/Composer2.5 Q4_K_M · Seconds / timed item 21.5s · Overall score 57.6 #25 #26 - Qwythos · Qwythos‑9B Claude Mythos Q8_0 · Seconds / timed item 21.1s · Overall score 52.9 #26 #27 - Ornith · Ornith‑1.0‑35B Q4_K_M · Seconds / timed item 118s · Overall score 50.1 #27 Seconds / timed item Overall score
Figure 4 Recorded tokens/item × cost Measured cost against recorded tokens / scored item. Each point is one tested model; hover or focus a point for its card.
Names stay on the top 10; every visible point keeps its rank label and hover detail.
Recorded tokens/item × cost Each point is one tested model. Recorded tokens / scored item is plotted on the X-axis and Measured cost is plotted on the Y-axis. 3.0k$0 8.2k$5.236 13.4k$10.473 18.6k$15.709 23.8k$20.945 #1 - DS-V4.1f · DeepSeek V4.1 Flash · Recorded tokens / scored item 10.5k · Measured cost $0.183 #1 #2 - Terra · GPT‑5.6 Terra · Recorded tokens / scored item 6.8k · Measured cost $1.867 #2 #3 - GPT5.5 · GPT‑5.5 · Recorded tokens / scored item 7.9k · Measured cost $4.822 #3 #4 - DS-V4f · DeepSeek V4 Flash · Recorded tokens / scored item 4.5k · Measured cost $0.290 #4 #5 - Sol · GPT‑5.6 Sol · Recorded tokens / scored item 6.9k · Measured cost $3.858 #5 #6 - Opus · Claude Opus 4.8 · Recorded tokens / scored item 12.0k · Measured cost $6.115 #6 #7 - Fable · Claude Fable 5 · Recorded tokens / scored item 13.1k · Measured cost $12.285 #7 #8 - Gemini · Gemini 3.5 Flash · Recorded tokens / scored item 8.4k · Measured cost $1.646 #8 #9 - GLM 5.3 Flash · GLM 5.3 Flash · Recorded tokens / scored item 8.6k · Measured cost $0.117 #9 #10 - Luna · GPT‑5.6 Luna · Recorded tokens / scored item 7.7k · Measured cost $1.000 #10 #11 - Opus 5.5 · Claude Opus 5.5 · Recorded tokens / scored item 13.0k · Measured cost $7.829 #11 #12 - Claude Fable 5.1 · Claude Fable 5.1 · Recorded tokens / scored item 12.9k · Measured cost $19.394 #12 #13 - GLM5.2 · GLM‑5.2 · Recorded tokens / scored item 22.4k · Measured cost $2.768 #13 #14 - Sonnet · Claude Sonnet 5 · Recorded tokens / scored item 13.7k · Measured cost $2.812 #14 #15 - DS-V4p · DeepSeek V4 Pro · Recorded tokens / scored item 9.6k · Measured cost $0.335 #15 #16 - GPT‑6 Astra · GPT‑6 Astra · Recorded tokens / scored item 6.8k · Measured cost $6.776 #16 #17 - Qwen3.7 · Qwen3.7 Max · Recorded tokens / scored item 8.7k · Measured cost $0.906 #17 #18 - M3 · MiniMax M3 · Recorded tokens / scored item 6.7k · Measured cost $0.182 #18 #19 - GPT‑6 Sol · GPT‑6 Sol · Recorded tokens / scored item 6.8k · Measured cost $1.310 #19 #20 - GPT‑6 Luna · GPT‑6 Luna · Recorded tokens / scored item 7.0k · Measured cost $0.070 #20 #21 - M3 Direct · MiniMax M3 Direct Plus · Recorded tokens / scored item 4.9k · Measured cost $0.169 #21 #22 - Kimi · Kimi K2.7 Code · Recorded tokens / scored item 8.7k · Measured cost $0.732 #22 #23 - Step · Step 3.7 Flash · Recorded tokens / scored item 11.5k · Measured cost $0.494 #23 #24 - Nemotron · NVIDIA Nemotron 3 Ultra · Recorded tokens / scored item 8.8k · Measured cost $0.489 #24 #25 - Gemma · Gemma4‑12B‑Coder Fable5/Composer2.5 Q4_K_M · Recorded tokens / scored item 8.3k · Measured cost $0 #25 #26 - Qwythos · Qwythos‑9B Claude Mythos Q8_0 · Recorded tokens / scored item 6.0k · Measured cost $0 #26 #27 - Ornith · Ornith‑1.0‑35B Q4_K_M · Recorded tokens / scored item 8.7k · Measured cost $0 #27 Recorded tokens / scored item Measured cost
Figure 5 Lane contrast Top eight entrants with Full, SWE, and Hard Intelligence shown side by side.
#1 DeepSeek V4.1 Flash
Full 89.83 SWE 89.52 Hard 88.11
#2 GPT‑5.6 Terra
Full 85.89 SWE 86.08 Hard 88.18
#3 GPT‑5.5
Full 84.33 SWE 82.53 Hard 88.28
#4 DeepSeek V4 Flash
Full 90.62 SWE 86.95 Hard 75.48
#5 GPT‑5.6 Sol
Full 84.29 SWE 76.12 Hard 91.31
#6 Claude Opus 4.8
Full 78.17 SWE 88.67 Hard 81.37
#7 Claude Fable 5
Full 72.05 SWE 81.42 Hard 92.78
#8 Gemini 3.5 Flash
Full 84.71 SWE 73.59 Hard 87.75

Hard Intelligence is shown as its own lane so cross-lane strengths and weaknesses stay visible.

Figure 6 Measured cost context Cost is shown because deployment economics matter, but it does not secretly rewrite capability scores.
DeepSeek V4.1 Flash $0.183 rank #1
GPT‑5.6 Terra $1.867 rank #2
GPT‑5.5 $4.822 rank #3
DeepSeek V4 Flash $0.290 rank #4
GPT‑5.6 Sol $3.858 rank #5
Claude Opus 4.8 $6.115 rank #6
Claude Fable 5 $12.285 rank #7
Gemini 3.5 Flash $1.646 rank #8
GLM 5.3 Flash $0.117 rank #9
GPT‑5.6 Luna $1.000 rank #10
Claude Opus 5.5 $7.829 rank #11
Claude Fable 5.1 $19.394 rank #12
GLM‑5.2 $2.768 rank #13
Claude Sonnet 5 $2.812 rank #14
DeepSeek V4 Pro $0.335 rank #15
GPT‑6 Astra $6.776 rank #16
Qwen3.7 Max $0.906 rank #17
MiniMax M3 $0.182 rank #18
GPT‑6 Sol $1.310 rank #19
GPT‑6 Luna $0.070 rank #20
MiniMax M3 Direct Plus $0.169 rank #21
Kimi K2.7 Code $0.732 rank #22
Step 3.7 Flash $0.494 rank #23
NVIDIA Nemotron 3 Ultra $0.489 rank #24

Very expensive rows are not punished twice; cost is visible telemetry and part of the public interpretation.

Figure 7 Lane balance pressure Largest gap between each entrant’s strongest and weakest measured major lane.
Step 3.7 Flash 53.10 Hard Intelligence 33.99 vs Full / Agentic 87.09 · rank #23
GPT‑6 Sol 40.03 SWE MVP 52.56 vs Hard Intelligence 92.59 · rank #19
GPT‑6 Astra 33.16 SWE MVP 61.92 vs Hard Intelligence 95.08 · rank #16
Claude Opus 5.5 31.12 SWE MVP 64.98 vs Hard Intelligence 96.10 · rank #11
Claude Fable 5.1 25.90 SWE MVP 69.92 vs Hard Intelligence 95.82 · rank #12
Qwythos‑9B Claude Mythos Q8_0 24.24 Hard Intelligence 43.91 vs Full / Agentic 68.15 · rank #26
GLM 5.3 Flash 24.10 SWE MVP 71.68 vs Hard Intelligence 95.78 · rank #9
Ornith‑1.0‑35B Q4_K_M 22.42 SWE MVP 39.16 vs Full / Agentic 61.58 · rank #27
GPT‑6 Luna 21.23 SWE MVP 62.59 vs Hard Intelligence 83.82 · rank #20
Gemma4‑12B‑Coder Fable5/Composer2.5 Q4_K_M 20.93 Hard Intelligence 48.96 vs Full / Agentic 69.89 · rank #25
Kimi K2.7 Code 20.86 SWE MVP 58.61 vs Full / Agentic 79.47 · rank #22
Claude Fable 5 20.73 Full / Agentic 72.05 vs Hard Intelligence 92.78 · rank #7
MiniMax M3 20.11 Hard Intelligence 66.77 vs SWE MVP 86.88 · rank #18
Qwen3.7 Max 19.09 Hard Intelligence 69.90 vs SWE MVP 88.99 · rank #17
GPT‑5.6 Sol 15.19 SWE MVP 76.12 vs Hard Intelligence 91.31 · rank #5
DeepSeek V4 Flash 15.14 Hard Intelligence 75.48 vs Full / Agentic 90.62 · rank #4
GLM‑5.2 15.00 Hard Intelligence 74.58 vs SWE MVP 89.58 · rank #13
Gemini 3.5 Flash 14.16 SWE MVP 73.59 vs Hard Intelligence 87.75 · rank #8
GPT‑5.6 Luna 11.10 SWE MVP 75.84 vs Full / Agentic 86.94 · rank #10
Claude Opus 4.8 10.50 Full / Agentic 78.17 vs SWE MVP 88.67 · rank #6
Claude Sonnet 5 9.81 Full / Agentic 76.72 vs Hard Intelligence 86.53 · rank #14
NVIDIA Nemotron 3 Ultra 8.52 SWE MVP 58.63 vs Full / Agentic 67.15 · rank #24
GPT‑5.5 5.75 SWE MVP 82.53 vs Hard Intelligence 88.28 · rank #3
MiniMax M3 Direct Plus 4.19 SWE MVP 67.77 vs Full / Agentic 71.96 · rank #21
GPT‑5.6 Terra 2.29 Full / Agentic 85.89 vs Hard Intelligence 88.18 · rank #2
DeepSeek V4.1 Flash 1.72 Hard Intelligence 88.11 vs Full / Agentic 89.83 · rank #1
DeepSeek V4 Pro 1.69 Hard Intelligence 78.38 vs Full / Agentic 80.06 · rank #15

Lower pressure means a more even profile; higher pressure explains why one strong lane may not lift the overall rank by itself.

Reading notes

Breadth wins the top spot

DeepSeek V4.1 Flash leads because its measured lanes stay high together: overall 89.15, Full 89.83, SWE 89.52, and Hard Intelligence 88.11.

Full / Agentic alone does not decide

DeepSeek V4 Flash owns Full rank #1 at 90.62, but the overall formula still checks SWE and Hard Intelligence before ordering the table.

SWE is a separate capability signal

GLM‑5.2 owns SWE rank #1 at 89.58. That lane rewards practical implementation and review behavior rather than only general prompt competence.

Hard Intelligence reshapes the table

Claude Opus 5.5 owns Hard Intelligence rank #1 at 96.10. That lane tests active inquiry, adaptation, repair, and authority integrity separately from Full and SWE.

The clearest drag is visible

Step 3.7 Flash has a Full/SWE average near 83.74, but Hard Intelligence is 33.99, so the blended overall lands at 67.16.

Table with reasons, not just numbers.

Each row states the score formula, lane ranks, cost context, and the main reason the entrant lands at its current position.

Ranking data
Rank Model Overall Full SWE Hard Intelligence Formula Cost and telemetry Why here
1 DeepSeek V4.1 FlashOpenCode Go relay · Full + SWE + Hard measured 89.15 89.83#2 89.52#2 88.11#10 mean(Full, SWE, Hard Intelligence) $0.18315.10s per itemtokens complete Overall 89.15 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full rank #2, SWE rank #2. Main limiter: Hard Intelligence at 88.11. Hard Intelligence contributes to the ranking as a separate measured lane.
2 GPT‑5.6 TerraOpenRouter · extra-high reasoning · Full + SWE + Hard measured 86.72 85.89#5 86.08#7 88.18#9 mean(Full, SWE, Hard Intelligence) $1.8675.75s per itemtokens complete Overall 86.72 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 88.18. Main limiter: Full / Agentic at 85.89. Hard Intelligence contributes to the ranking as a separate measured lane.
3 GPT‑5.5ChatGPT Codex subscription · gpt-5.5 · extra-high reasoning · Full + SWE + Hard measured 85.05 84.33#7 82.53#8 88.28#8 mean(Full, SWE, Hard Intelligence) $4.82219.59s per itemtokens complete Overall 85.05 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 88.28. Main limiter: SWE MVP at 82.53. Hard Intelligence contributes to the ranking as a separate measured lane.
4 DeepSeek V4 FlashDeepSeek direct API · refreshed Hard Intelligence telemetry 84.35 90.62#1 86.95#5 75.48#17 mean(Full, SWE, Hard Intelligence) $0.29022.93s per itemtokens complete Overall 84.35 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full rank #1. Main limiter: Hard Intelligence at 75.48. Hard Intelligence contributes to the ranking as a separate measured lane.
5 GPT‑5.6 SolOpenRouter · extra-high reasoning · Full + SWE + Hard measured 83.91 84.29#8 76.12#13 91.31#7 mean(Full, SWE, Hard Intelligence) $3.85813.30s per itemtokens complete Overall 83.91 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 91.31. Main limiter: SWE MVP at 76.12. Hard Intelligence contributes to the ranking as a separate measured lane.
6 Claude Opus 4.8OpenRouter · extra-high reasoning · Hard Intelligence 82.74 78.17#14 88.67#4 81.37#14 mean(Full, SWE, Hard Intelligence) $6.11510.43s per itemtokens complete Overall 82.74 uses mean(Full, SWE, Hard Intelligence). Strength signal: SWE MVP at 88.67. Main limiter: Full / Agentic at 78.17. Hard Intelligence contributes to the ranking as a separate measured lane.
7 Claude Fable 5OpenRouter · extra-high reasoning · Full + SWE + Hard measured 82.08 72.05#22 81.42#9 92.78#5 mean(Full, SWE, Hard Intelligence) $12.28512.37s per itemtokens complete Overall 82.08 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 92.78. Main limiter: Full / Agentic at 72.05. Hard Intelligence contributes to the ranking as a separate measured lane.
8 Gemini 3.5 FlashOpenRouter · extra-high reasoning · Hard Intelligence 82.02 84.71#6 73.59#15 87.75#11 mean(Full, SWE, Hard Intelligence) $1.6469.13s per itemtokens complete Overall 82.02 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 87.75. Main limiter: SWE MVP at 73.59. Hard Intelligence contributes to the ranking as a separate measured lane.
9 GLM 5.3 Flashz.ai Coding Plan · glm-5.3-flash · maximum reasoning with thinking enabled · Full + SWE + Hard measured 81.32 76.49#19 71.68#16 95.78#3 mean(Full, SWE, Hard Intelligence) $0.11735.45s per itemtokens complete Overall 81.32 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence rank #3. Main limiter: SWE MVP at 71.68. Hard Intelligence contributes to the ranking as a separate measured lane.
10 GPT‑5.6 LunaOpenRouter · extra-high reasoning · Full + SWE + Hard measured 81.15 86.94#4 75.84#14 80.66#15 mean(Full, SWE, Hard Intelligence) $1.0008.02s per itemtokens complete Overall 81.15 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 86.94. Main limiter: SWE MVP at 75.84. Hard Intelligence contributes to the ranking as a separate measured lane.
11 Claude Opus 5.5Claude Max subscription · extra-high reasoning · Full + SWE + Hard measured 80.87 81.54#10 64.98#19 96.10#1 mean(Full, SWE, Hard Intelligence) $7.8297.42s per itemtokens complete Overall 80.87 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence rank #1. Main limiter: SWE MVP at 64.98. Hard Intelligence contributes to the ranking as a separate measured lane.
12 Claude Fable 5.1Claude Max subscription · extra-high reasoning · Full + SWE + Hard measured 80.53 75.84#20 69.92#17 95.82#2 mean(Full, SWE, Hard Intelligence) $19.3948.97s per itemtokens complete Overall 80.53 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence rank #2. Main limiter: SWE MVP at 69.92. Hard Intelligence contributes to the ranking as a separate measured lane.
13 GLM‑5.2OpenRouter · z-ai/glm-5.2 · maximum reasoning · Full + SWE + Hard measured 80.50 77.35#17 89.58#1 74.58#18 mean(Full, SWE, Hard Intelligence) $2.76891.19s per itemtokens complete Overall 80.50 uses mean(Full, SWE, Hard Intelligence). Strength signal: SWE rank #1. Main limiter: Hard Intelligence at 74.58. Hard Intelligence contributes to the ranking as a separate measured lane.
14 Claude Sonnet 5OpenRouter · extra-high reasoning · Full + SWE + Hard measured 80.44 76.72#18 78.08#12 86.53#12 mean(Full, SWE, Hard Intelligence) $2.81212.51s per itemtokens complete Overall 80.44 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 86.53. Main limiter: Full / Agentic at 76.72. Hard Intelligence contributes to the ranking as a separate measured lane.
15 DeepSeek V4 ProDeepSeek direct API · maximum-reasoning Hard IQ 79.37 80.06#11 79.68#11 78.38#16 mean(Full, SWE, Hard Intelligence) $0.33523.25s per itemtokens complete Overall 79.37 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 80.06. Main limiter: Hard Intelligence at 78.38. Hard Intelligence contributes to the ranking as a separate measured lane.
16 GPT‑6 AstraChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured 78.29 77.86#16 61.92#21 95.08#4 mean(Full, SWE, Hard Intelligence) $6.77613.48s per itemtokens complete Overall 78.29 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 95.08. Main limiter: SWE MVP at 61.92. Hard Intelligence contributes to the ranking as a separate measured lane.
17 Qwen3.7 MaxOpenRouter · extra-high reasoning · Hard Intelligence 77.52 73.67#21 88.99#3 69.90#20 mean(Full, SWE, Hard Intelligence) $0.90622.71s per itemtokens complete Overall 77.52 uses mean(Full, SWE, Hard Intelligence). Strength signal: SWE rank #3. Main limiter: Hard Intelligence at 69.90. Hard Intelligence contributes to the ranking as a separate measured lane.
18 MiniMax M3OpenRouter · extra-high reasoning 77.19 77.92#15 86.88#6 66.77#21 mean(Full, SWE, Hard Intelligence) $0.18227.52s per itemtokens complete Overall 77.19 uses mean(Full, SWE, Hard Intelligence). Strength signal: SWE MVP at 86.88. Main limiter: Hard Intelligence at 66.77. Hard Intelligence contributes to the ranking as a separate measured lane.
19 GPT‑6 SolChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured 76.09 83.13#9 52.56#25 92.59#6 mean(Full, SWE, Hard Intelligence) $1.31012.11s per itemtokens complete Overall 76.09 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 92.59. Main limiter: SWE MVP at 52.56. Hard Intelligence contributes to the ranking as a separate measured lane.
20 GPT‑6 LunaChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured 74.99 78.57#13 62.59#20 83.82#13 mean(Full, SWE, Hard Intelligence) $0.07013.43s per itemtokens complete Overall 74.99 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 83.82. Main limiter: SWE MVP at 62.59. Hard Intelligence contributes to the ranking as a separate measured lane.
21 MiniMax M3 Direct PlusMiniMax direct API · extra-high reasoning · costed at MiniMax's published M3 API rates 70.51 71.96#23 67.77#18 71.80#19 mean(Full, SWE, Hard Intelligence) $0.16921.94s per itemtokens complete Overall 70.51 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 71.96. Main limiter: SWE MVP at 67.77. Hard Intelligence contributes to the ranking as a separate measured lane.
22 Kimi K2.7 CodeOpenRouter · Kimi K2.7 Code · extra-high Hard IQ 68.18 79.47#12 58.61#23 66.46#22 mean(Full, SWE, Hard Intelligence) $0.73228.60s per itemtokens complete Overall 68.18 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 79.47. Main limiter: SWE MVP at 58.61. Hard Intelligence contributes to the ranking as a separate measured lane.
23 Step 3.7 FlashOpenRouter · stepfun/step-3.7-flash · extra-high reasoning · Full + SWE + Hard measured 67.16 87.09#3 80.39#10 33.99#27 mean(Full, SWE, Hard Intelligence) $0.49427.77s per itemtokens complete Overall 67.16 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full rank #3. Main limiter: Hard Intelligence at 33.99. Hard Intelligence contributes to the ranking as a separate measured lane.
24 NVIDIA Nemotron 3 UltraOpenRouter · nvidia/nemotron-3-ultra-550b-a55b · extra-high reasoning 63.45 67.15#26 58.63#22 64.56#23 mean(Full, SWE, Hard Intelligence) $0.48970.71s per itemtokens complete Overall 63.45 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 67.15. Main limiter: SWE MVP at 58.63. Hard Intelligence contributes to the ranking as a separate measured lane.
25 Gemma4‑12B‑Coder Fable5/Composer2.5 Q4_K_MLocal modelLocal GGUF · llama.cpp Vulkan · Q4_K_M 57.62 69.89#24 54.01#24 48.96#25 mean(Full, SWE, Hard Intelligence) $021.52s per itemtokens complete Overall 57.62 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 69.89. Main limiter: Hard Intelligence at 48.96. Hard Intelligence contributes to the ranking as a separate measured lane.
26 Qwythos‑9B Claude Mythos Q8_0Local modelLocal GGUF · llama.cpp Vulkan · Q8_0 · 256K allocation verified 52.86 68.15#25 46.51#26 43.91#26 mean(Full, SWE, Hard Intelligence) $021.13s per itemtokens complete Overall 52.86 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 68.15. Main limiter: Hard Intelligence at 43.91. Hard Intelligence contributes to the ranking as a separate measured lane.
27 Ornith‑1.0‑35B Q4_K_MLocal modelLocal GGUF · llama.cpp Vulkan · Q4_K_M · 35B MoE 50.11 61.58#27 39.16#27 49.58#24 mean(Full, SWE, Hard Intelligence) $0117.66s per itemtokens complete Overall 50.11 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 61.58. Main limiter: SWE MVP at 39.16. Hard Intelligence contributes to the ranking as a separate measured lane.
Interpretation

Why the leader leads

LeaderDeepSeek V4.1 Flash
Overall89.15
Full89.83
SWE89.52
Hard IQ88.11

The top rank belongs to the entrant with the strongest cross-lane balance under the current formula, not simply the best isolated lane score.

Lane policy

How Hard Intelligence is handled

Scopeactive inquiry + adaptation + repair
Formula roleincluded when measured
Blank cellsnot yet measured
Interpretationseparate from Full and SWE

When a Hard Intelligence score is published, it becomes the third major lane in the overall mean. Otherwise the row remains ranked by the measured lanes it has.

Tie-break reading

How to compare close rows

Overallfirst glance
Lane ranksdiagnosis
Costruntime context
Reliabilityoperational risk

Close overall scores should be read through the lane breakdown. A model can be strong for building software while weaker at active inquiry, or the reverse.