Agentic discipline
Full / Agentic laneStructured outputs, tool-use boundaries, instruction following, grounded reasoning, and hallucination resistance.
Independent model evaluation from Resyst Labs
Every ranked model is measured on agentic discipline, software engineering, and Hard Intelligence, and every score is published beside the data that produced it. Resyst Arena adds turn-by-turn duels you can replay.
Three scored lanes and one tactical testbed. Each strip plots every ranked model on that lane; the brightest mark is the lane leader.
Structured outputs, tool-use boundaries, instruction following, grounded reasoning, and hallucination resistance.
Practical implementation quality, final-answer usefulness, source handling, and architecture cleanliness.
Active inquiry, online adaptation, evidence-driven self-repair, and authority integrity under a public hard-reasoning diagnostic.
Turn-based spatial duels where legal action discipline and tactical continuity are measured separately from runtime telemetry.
Local and API-backed systems share a single tournament view. Provider, runtime, cost, reliability, and lane basis remain visible so comparisons stay honest.
OpenCode Go relay · Full + SWE + Hard measured
OpenRouter · extra-high reasoning · Full + SWE + Hard measured
ChatGPT Codex subscription · gpt-5.5 · extra-high reasoning · Full + SWE + Hard measured
| Rank | Model | Overall | Full | SWE | Hard Intelligence | Cost | Reliability | Result page |
|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4.1 Flash OpenCode Go relay · Full + SWE + Hard measured | 89.15 | 89.83#2 | 89.52#2 | 88.11#10 | $0.183 | 100.0% | Result |
| 2 | GPT‑5.6 Terra OpenRouter · extra-high reasoning · Full + SWE + Hard measured | 86.72 | 85.89#5 | 86.08#7 | 88.18#9 | $1.867 | 100.0% | Result |
| 3 | GPT‑5.5 ChatGPT Codex subscription · gpt-5.5 · extra-high reasoning · Full + SWE + Hard measured | 85.05 | 84.33#7 | 82.53#8 | 88.28#8 | $4.822 | 100.0% | Result |
| 4 | DeepSeek V4 Flash DeepSeek direct API · refreshed Hard Intelligence telemetry | 84.35 | 90.62#1 | 86.95#5 | 75.48#17 | $0.290 | 100.0% | Result |
| 5 | GPT‑5.6 Sol OpenRouter · extra-high reasoning · Full + SWE + Hard measured | 83.91 | 84.29#8 | 76.12#13 | 91.31#7 | $3.858 | 100.0% | Result |
| 6 | Claude Opus 4.8 OpenRouter · extra-high reasoning · Hard Intelligence | 82.74 | 78.17#14 | 88.67#4 | 81.37#14 | $6.115 | 100.0% | Result |
| 7 | Claude Fable 5 OpenRouter · extra-high reasoning · Full + SWE + Hard measured | 82.08 | 72.05#22 | 81.42#9 | 92.78#5 | $12.285 | 100.0% | Result |
| 8 | Gemini 3.5 Flash OpenRouter · extra-high reasoning · Hard Intelligence | 82.02 | 84.71#6 | 73.59#15 | 87.75#11 | $1.646 | 100.0% | Result |
| 9 | GLM 5.3 Flash z.ai Coding Plan · glm-5.3-flash · maximum reasoning with thinking enabled · Full + SWE + Hard measured | 81.32 | 76.49#19 | 71.68#16 | 95.78#3 | $0.117 | 99.2% | Result |
| 10 | GPT‑5.6 Luna OpenRouter · extra-high reasoning · Full + SWE + Hard measured | 81.15 | 86.94#4 | 75.84#14 | 80.66#15 | $1.000 | 100.0% | Result |
| 11 | Claude Opus 5.5 Claude Max subscription · extra-high reasoning · Full + SWE + Hard measured | 80.87 | 81.54#10 | 64.98#19 | 96.10#1 | $7.829 | 100.0% | Result |
| 12 | Claude Fable 5.1 Claude Max subscription · extra-high reasoning · Full + SWE + Hard measured | 80.53 | 75.84#20 | 69.92#17 | 95.82#2 | $19.394 | 100.0% | Result |
| 13 | GLM‑5.2 OpenRouter · z-ai/glm-5.2 · maximum reasoning · Full + SWE + Hard measured | 80.50 | 77.35#17 | 89.58#1 | 74.58#18 | $2.768 | 100.0% | Result |
| 14 | Claude Sonnet 5 OpenRouter · extra-high reasoning · Full + SWE + Hard measured | 80.44 | 76.72#18 | 78.08#12 | 86.53#12 | $2.812 | 100.0% | Result |
| 15 | DeepSeek V4 Pro DeepSeek direct API · maximum-reasoning Hard IQ | 79.37 | 80.06#11 | 79.68#11 | 78.38#16 | $0.335 | 100.0% | Result |
| 16 | GPT‑6 Astra ChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured | 78.29 | 77.86#16 | 61.92#21 | 95.08#4 | $6.776 | 100.0% | Result |
| 17 | Qwen3.7 Max OpenRouter · extra-high reasoning · Hard Intelligence | 77.52 | 73.67#21 | 88.99#3 | 69.90#20 | $0.906 | 100.0% | Result |
| 18 | MiniMax M3 OpenRouter · extra-high reasoning | 77.19 | 77.92#15 | 86.88#6 | 66.77#21 | $0.182 | 100.0% | Result |
| 19 | GPT‑6 Sol ChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured | 76.09 | 83.13#9 | 52.56#25 | 92.59#6 | $1.310 | 100.0% | Result |
| 20 | GPT‑6 Luna ChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured | 74.99 | 78.57#13 | 62.59#20 | 83.82#13 | $0.070 | 100.0% | Result |
| 21 | MiniMax M3 Direct Plus MiniMax direct API · extra-high reasoning · costed at MiniMax's published M3 API rates | 70.51 | 71.96#23 | 67.77#18 | 71.80#19 | $0.169 | 100.0% | Result |
| 22 | Kimi K2.7 Code OpenRouter · Kimi K2.7 Code · extra-high Hard IQ | 68.18 | 79.47#12 | 58.61#23 | 66.46#22 | $0.732 | 100.0% | Result |
| 23 | Step 3.7 Flash OpenRouter · stepfun/step-3.7-flash · extra-high reasoning · Full + SWE + Hard measured | 67.16 | 87.09#3 | 80.39#10 | 33.99#27 | $0.494 | 100.0% | Result |
| 24 | NVIDIA Nemotron 3 Ultra OpenRouter · nvidia/nemotron-3-ultra-550b-a55b · extra-high reasoning | 63.45 | 67.15#26 | 58.63#22 | 64.56#23 | $0.489 | 100.0% | Result |
| 25 | Gemma4‑12B‑Coder Fable5/Composer2.5 Q4_K_M Local model Local GGUF · llama.cpp Vulkan · Q4_K_M | 57.62 | 69.89#24 | 54.01#24 | 48.96#25 | $0 | 100.0% | Result |
| 26 | Qwythos‑9B Claude Mythos Q8_0 Local model Local GGUF · llama.cpp Vulkan · Q8_0 · 256K allocation verified | 52.86 | 68.15#25 | 46.51#26 | 43.91#26 | $0 | 100.0% | Result |
| 27 | Ornith‑1.0‑35B Q4_K_M Local model Local GGUF · llama.cpp Vulkan · Q4_K_M · 35B MoE | 50.11 | 61.58#27 | 39.16#27 | 49.58#24 | $0 | 100.0% | Result |
Resyst Arena evaluates spatial strategy in deterministic turn-based games. Each public match summary links to a replay with board states, legal actions, events, and tactical telemetry.
Kimi K2.7 Code leads 2–0. Replays stay grouped under the model-vs-model encounter, including side-swapped rounds.
DeepSeek V4 Flash leads 2–0. Replays stay grouped under the model-vs-model encounter, including side-swapped rounds.
Gemini 3 Flash Preview won the replay. Replays stay grouped under the model-vs-model encounter, including side-swapped rounds.
Agentic, software-engineering, and Hard Intelligence diagnostics are preserved as distinct measurements before any publication formula combines them.
The same model can appear through different providers or runtimes. The table exposes basis metadata instead of hiding infrastructure differences.
Single runs are evidence records. Stronger claims require repeated series, side swaps, seed variation, and comparable scoring settings.
Benchmark summaries are published as versioned data files. The presentation layer is intentionally separate from the scoring harness, so rankings can evolve without rewriting the public record.