27 with Hard Intelligence data
Unified ranking, lane-aware explanation
Why the ranking looks like this.
The public ranking is not a single vibe score. It orders measured entrants by a transparent overall formula while keeping Full / Agentic, SWE MVP, Hard Intelligence, cost, and reliability visible.
Overall 89.15
#1 to #27
Full + SWE + published Hard Intelligence when measured
Static HTML plus public JSON
Overall is a lane mean, not a hidden replacement for source measurements.
Tradeoff scatter maps
Each point is one tested model at the intersection of two public telemetry axes. Use the maps to read quality against cost, speed, and recorded token use. Runtime and token axes are normalized per item and show coverage in the hover cards; they are telemetry, not current overall score inputs.
Hard Intelligence is shown as its own lane so cross-lane strengths and weaknesses stay visible.
Very expensive rows are not punished twice; cost is visible telemetry and part of the public interpretation.
Lower pressure means a more even profile; higher pressure explains why one strong lane may not lift the overall rank by itself.
Reading notes
Breadth wins the top spot
DeepSeek V4.1 Flash leads because its measured lanes stay high together: overall 89.15, Full 89.83, SWE 89.52, and Hard Intelligence 88.11.
Full / Agentic alone does not decide
DeepSeek V4 Flash owns Full rank #1 at 90.62, but the overall formula still checks SWE and Hard Intelligence before ordering the table.
SWE is a separate capability signal
GLM‑5.2 owns SWE rank #1 at 89.58. That lane rewards practical implementation and review behavior rather than only general prompt competence.
Hard Intelligence reshapes the table
Claude Opus 5.5 owns Hard Intelligence rank #1 at 96.10. That lane tests active inquiry, adaptation, repair, and authority integrity separately from Full and SWE.
The clearest drag is visible
Step 3.7 Flash has a Full/SWE average near 83.74, but Hard Intelligence is 33.99, so the blended overall lands at 67.16.
Table with reasons, not just numbers.
Each row states the score formula, lane ranks, cost context, and the main reason the entrant lands at its current position.
| Rank | Model | Overall | Full | SWE | Hard Intelligence | Formula | Cost and telemetry | Why here |
|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4.1 FlashOpenCode Go relay · Full + SWE + Hard measured | 89.15 | 89.83#2 | 89.52#2 | 88.11#10 | mean(Full, SWE, Hard Intelligence) | $0.18315.10s per itemtokens complete | Overall 89.15 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full rank #2, SWE rank #2. Main limiter: Hard Intelligence at 88.11. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 2 | GPT‑5.6 TerraOpenRouter · extra-high reasoning · Full + SWE + Hard measured | 86.72 | 85.89#5 | 86.08#7 | 88.18#9 | mean(Full, SWE, Hard Intelligence) | $1.8675.75s per itemtokens complete | Overall 86.72 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 88.18. Main limiter: Full / Agentic at 85.89. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 3 | GPT‑5.5ChatGPT Codex subscription · gpt-5.5 · extra-high reasoning · Full + SWE + Hard measured | 85.05 | 84.33#7 | 82.53#8 | 88.28#8 | mean(Full, SWE, Hard Intelligence) | $4.82219.59s per itemtokens complete | Overall 85.05 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 88.28. Main limiter: SWE MVP at 82.53. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 4 | DeepSeek V4 FlashDeepSeek direct API · refreshed Hard Intelligence telemetry | 84.35 | 90.62#1 | 86.95#5 | 75.48#17 | mean(Full, SWE, Hard Intelligence) | $0.29022.93s per itemtokens complete | Overall 84.35 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full rank #1. Main limiter: Hard Intelligence at 75.48. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 5 | GPT‑5.6 SolOpenRouter · extra-high reasoning · Full + SWE + Hard measured | 83.91 | 84.29#8 | 76.12#13 | 91.31#7 | mean(Full, SWE, Hard Intelligence) | $3.85813.30s per itemtokens complete | Overall 83.91 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 91.31. Main limiter: SWE MVP at 76.12. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 6 | Claude Opus 4.8OpenRouter · extra-high reasoning · Hard Intelligence | 82.74 | 78.17#14 | 88.67#4 | 81.37#14 | mean(Full, SWE, Hard Intelligence) | $6.11510.43s per itemtokens complete | Overall 82.74 uses mean(Full, SWE, Hard Intelligence). Strength signal: SWE MVP at 88.67. Main limiter: Full / Agentic at 78.17. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 7 | Claude Fable 5OpenRouter · extra-high reasoning · Full + SWE + Hard measured | 82.08 | 72.05#22 | 81.42#9 | 92.78#5 | mean(Full, SWE, Hard Intelligence) | $12.28512.37s per itemtokens complete | Overall 82.08 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 92.78. Main limiter: Full / Agentic at 72.05. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 8 | Gemini 3.5 FlashOpenRouter · extra-high reasoning · Hard Intelligence | 82.02 | 84.71#6 | 73.59#15 | 87.75#11 | mean(Full, SWE, Hard Intelligence) | $1.6469.13s per itemtokens complete | Overall 82.02 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 87.75. Main limiter: SWE MVP at 73.59. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 9 | GLM 5.3 Flashz.ai Coding Plan · glm-5.3-flash · maximum reasoning with thinking enabled · Full + SWE + Hard measured | 81.32 | 76.49#19 | 71.68#16 | 95.78#3 | mean(Full, SWE, Hard Intelligence) | $0.11735.45s per itemtokens complete | Overall 81.32 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence rank #3. Main limiter: SWE MVP at 71.68. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 10 | GPT‑5.6 LunaOpenRouter · extra-high reasoning · Full + SWE + Hard measured | 81.15 | 86.94#4 | 75.84#14 | 80.66#15 | mean(Full, SWE, Hard Intelligence) | $1.0008.02s per itemtokens complete | Overall 81.15 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 86.94. Main limiter: SWE MVP at 75.84. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 11 | Claude Opus 5.5Claude Max subscription · extra-high reasoning · Full + SWE + Hard measured | 80.87 | 81.54#10 | 64.98#19 | 96.10#1 | mean(Full, SWE, Hard Intelligence) | $7.8297.42s per itemtokens complete | Overall 80.87 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence rank #1. Main limiter: SWE MVP at 64.98. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 12 | Claude Fable 5.1Claude Max subscription · extra-high reasoning · Full + SWE + Hard measured | 80.53 | 75.84#20 | 69.92#17 | 95.82#2 | mean(Full, SWE, Hard Intelligence) | $19.3948.97s per itemtokens complete | Overall 80.53 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence rank #2. Main limiter: SWE MVP at 69.92. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 13 | GLM‑5.2OpenRouter · z-ai/glm-5.2 · maximum reasoning · Full + SWE + Hard measured | 80.50 | 77.35#17 | 89.58#1 | 74.58#18 | mean(Full, SWE, Hard Intelligence) | $2.76891.19s per itemtokens complete | Overall 80.50 uses mean(Full, SWE, Hard Intelligence). Strength signal: SWE rank #1. Main limiter: Hard Intelligence at 74.58. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 14 | Claude Sonnet 5OpenRouter · extra-high reasoning · Full + SWE + Hard measured | 80.44 | 76.72#18 | 78.08#12 | 86.53#12 | mean(Full, SWE, Hard Intelligence) | $2.81212.51s per itemtokens complete | Overall 80.44 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 86.53. Main limiter: Full / Agentic at 76.72. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 15 | DeepSeek V4 ProDeepSeek direct API · maximum-reasoning Hard IQ | 79.37 | 80.06#11 | 79.68#11 | 78.38#16 | mean(Full, SWE, Hard Intelligence) | $0.33523.25s per itemtokens complete | Overall 79.37 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 80.06. Main limiter: Hard Intelligence at 78.38. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 16 | GPT‑6 AstraChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured | 78.29 | 77.86#16 | 61.92#21 | 95.08#4 | mean(Full, SWE, Hard Intelligence) | $6.77613.48s per itemtokens complete | Overall 78.29 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 95.08. Main limiter: SWE MVP at 61.92. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 17 | Qwen3.7 MaxOpenRouter · extra-high reasoning · Hard Intelligence | 77.52 | 73.67#21 | 88.99#3 | 69.90#20 | mean(Full, SWE, Hard Intelligence) | $0.90622.71s per itemtokens complete | Overall 77.52 uses mean(Full, SWE, Hard Intelligence). Strength signal: SWE rank #3. Main limiter: Hard Intelligence at 69.90. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 18 | MiniMax M3OpenRouter · extra-high reasoning | 77.19 | 77.92#15 | 86.88#6 | 66.77#21 | mean(Full, SWE, Hard Intelligence) | $0.18227.52s per itemtokens complete | Overall 77.19 uses mean(Full, SWE, Hard Intelligence). Strength signal: SWE MVP at 86.88. Main limiter: Hard Intelligence at 66.77. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 19 | GPT‑6 SolChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured | 76.09 | 83.13#9 | 52.56#25 | 92.59#6 | mean(Full, SWE, Hard Intelligence) | $1.31012.11s per itemtokens complete | Overall 76.09 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 92.59. Main limiter: SWE MVP at 52.56. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 20 | GPT‑6 LunaChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured | 74.99 | 78.57#13 | 62.59#20 | 83.82#13 | mean(Full, SWE, Hard Intelligence) | $0.07013.43s per itemtokens complete | Overall 74.99 uses mean(Full, SWE, Hard Intelligence). Strength signal: Hard Intelligence at 83.82. Main limiter: SWE MVP at 62.59. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 21 | MiniMax M3 Direct PlusMiniMax direct API · extra-high reasoning · costed at MiniMax's published M3 API rates | 70.51 | 71.96#23 | 67.77#18 | 71.80#19 | mean(Full, SWE, Hard Intelligence) | $0.16921.94s per itemtokens complete | Overall 70.51 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 71.96. Main limiter: SWE MVP at 67.77. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 22 | Kimi K2.7 CodeOpenRouter · Kimi K2.7 Code · extra-high Hard IQ | 68.18 | 79.47#12 | 58.61#23 | 66.46#22 | mean(Full, SWE, Hard Intelligence) | $0.73228.60s per itemtokens complete | Overall 68.18 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 79.47. Main limiter: SWE MVP at 58.61. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 23 | Step 3.7 FlashOpenRouter · stepfun/step-3.7-flash · extra-high reasoning · Full + SWE + Hard measured | 67.16 | 87.09#3 | 80.39#10 | 33.99#27 | mean(Full, SWE, Hard Intelligence) | $0.49427.77s per itemtokens complete | Overall 67.16 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full rank #3. Main limiter: Hard Intelligence at 33.99. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 24 | NVIDIA Nemotron 3 UltraOpenRouter · nvidia/nemotron-3-ultra-550b-a55b · extra-high reasoning | 63.45 | 67.15#26 | 58.63#22 | 64.56#23 | mean(Full, SWE, Hard Intelligence) | $0.48970.71s per itemtokens complete | Overall 63.45 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 67.15. Main limiter: SWE MVP at 58.63. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 25 | Gemma4‑12B‑Coder Fable5/Composer2.5 Q4_K_MLocal modelLocal GGUF · llama.cpp Vulkan · Q4_K_M | 57.62 | 69.89#24 | 54.01#24 | 48.96#25 | mean(Full, SWE, Hard Intelligence) | $021.52s per itemtokens complete | Overall 57.62 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 69.89. Main limiter: Hard Intelligence at 48.96. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 26 | Qwythos‑9B Claude Mythos Q8_0Local modelLocal GGUF · llama.cpp Vulkan · Q8_0 · 256K allocation verified | 52.86 | 68.15#25 | 46.51#26 | 43.91#26 | mean(Full, SWE, Hard Intelligence) | $021.13s per itemtokens complete | Overall 52.86 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 68.15. Main limiter: Hard Intelligence at 43.91. Hard Intelligence contributes to the ranking as a separate measured lane. |
| 27 | Ornith‑1.0‑35B Q4_K_MLocal modelLocal GGUF · llama.cpp Vulkan · Q4_K_M · 35B MoE | 50.11 | 61.58#27 | 39.16#27 | 49.58#24 | mean(Full, SWE, Hard Intelligence) | $0117.66s per itemtokens complete | Overall 50.11 uses mean(Full, SWE, Hard Intelligence). Strength signal: Full / Agentic at 61.58. Main limiter: SWE MVP at 39.16. Hard Intelligence contributes to the ranking as a separate measured lane. |
Why the leader leads
The top rank belongs to the entrant with the strongest cross-lane balance under the current formula, not simply the best isolated lane score.
How Hard Intelligence is handled
When a Hard Intelligence score is published, it becomes the third major lane in the overall mean. Otherwise the row remains ranked by the measured lanes it has.
How to compare close rows
Close overall scores should be read through the lane breakdown. A model can be strong for building software while weaker at active inquiry, or the reverse.