Independent model evaluation from Resyst Labs

AI model benchmarks, with receipts.

Every ranked model is measured on agentic discipline, software engineering, and Hard Intelligence, and every score is published beside the data that produced it. Resyst Arena adds turn-by-turn duels you can replay.

Figure 1 Overall score spectrum Every ranked model at its overall score. Gold marks API-backed systems, cyan marks local hardware. Hover or focus a mark for the model.
  1. 1DeepSeek V4.1 Flash89.15
  2. 2GPT‑5.6 Terra86.72
  3. 3GPT‑5.585.05
27
ranked models
3
scored lanes
5
Arena replays
Sep 25, 2026
data refresh

What is measured

Three scored lanes and one tactical testbed. Each strip plots every ranked model on that lane; the brightest mark is the lane leader.

Agentic discipline

Full / Agentic lane

Structured outputs, tool-use boundaries, instruction following, grounded reasoning, and hallucination resistance.

27 models measured, from 61.6 to 90.6.
Lane leader DeepSeek V4 Flash 90.6

Software execution

SWE MVP lane

Practical implementation quality, final-answer usefulness, source handling, and architecture cleanliness.

27 models measured, from 39.2 to 89.6.
Lane leader GLM‑5.2 89.6

Hard Intelligence

Hard Intelligence lane

Active inquiry, online adaptation, evidence-driven self-repair, and authority integrity under a public hard-reasoning diagnostic.

27 models measured, from 34.0 to 96.1.
Lane leader Claude Opus 5.5 96.1

Resyst Arena

Tactical testbed

Turn-based spatial duels where legal action discipline and tactical continuity are measured separately from runtime telemetry.

5 replays 3 encounters 483 recorded turns
Replay room Open Resyst Arena

One table, visible tradeoffs.

Local and API-backed systems share a single tournament view. Provider, runtime, cost, reliability, and lane basis remain visible so comparisons stay honest.

1

DeepSeek V4.1 Flash

OpenCode Go relay · Full + SWE + Hard measured

89.15overall
Full
89.83
SWE
89.52
Hard Intelligence
88.11
Open result
2

GPT‑5.6 Terra

OpenRouter · extra-high reasoning · Full + SWE + Hard measured

86.72overall
Full
85.89
SWE
86.08
Hard Intelligence
88.18
Open result
3

GPT‑5.5

ChatGPT Codex subscription · gpt-5.5 · extra-high reasoning · Full + SWE + Hard measured

85.05overall
Full
84.33
SWE
82.53
Hard Intelligence
88.28
Open result
Figure 2 Unified ranking Overall is the mean of the measured lanes. Bars are drawn to the same 0 to 100 scale in every column. Hard Intelligence shows its four sub-lanes beneath the score.
Rank Model Overall Full SWE Hard Intelligence Cost Reliability Result page
1 DeepSeek V4.1 Flash OpenCode Go relay · Full + SWE + Hard measured 89.15 89.83#2 89.52#2 88.11#10 $0.183 100.0% Result
2 GPT‑5.6 Terra OpenRouter · extra-high reasoning · Full + SWE + Hard measured 86.72 85.89#5 86.08#7 88.18#9 $1.867 100.0% Result
3 GPT‑5.5 ChatGPT Codex subscription · gpt-5.5 · extra-high reasoning · Full + SWE + Hard measured 85.05 84.33#7 82.53#8 88.28#8 $4.822 100.0% Result
4 DeepSeek V4 Flash DeepSeek direct API · refreshed Hard Intelligence telemetry 84.35 90.62#1 86.95#5 75.48#17 $0.290 100.0% Result
5 GPT‑5.6 Sol OpenRouter · extra-high reasoning · Full + SWE + Hard measured 83.91 84.29#8 76.12#13 91.31#7 $3.858 100.0% Result
6 Claude Opus 4.8 OpenRouter · extra-high reasoning · Hard Intelligence 82.74 78.17#14 88.67#4 81.37#14 $6.115 100.0% Result
7 Claude Fable 5 OpenRouter · extra-high reasoning · Full + SWE + Hard measured 82.08 72.05#22 81.42#9 92.78#5 $12.285 100.0% Result
8 Gemini 3.5 Flash OpenRouter · extra-high reasoning · Hard Intelligence 82.02 84.71#6 73.59#15 87.75#11 $1.646 100.0% Result
9 GLM 5.3 Flash z.ai Coding Plan · glm-5.3-flash · maximum reasoning with thinking enabled · Full + SWE + Hard measured 81.32 76.49#19 71.68#16 95.78#3 $0.117 99.2% Result
10 GPT‑5.6 Luna OpenRouter · extra-high reasoning · Full + SWE + Hard measured 81.15 86.94#4 75.84#14 80.66#15 $1.000 100.0% Result
11 Claude Opus 5.5 Claude Max subscription · extra-high reasoning · Full + SWE + Hard measured 80.87 81.54#10 64.98#19 96.10#1 $7.829 100.0% Result
12 Claude Fable 5.1 Claude Max subscription · extra-high reasoning · Full + SWE + Hard measured 80.53 75.84#20 69.92#17 95.82#2 $19.394 100.0% Result
13 GLM‑5.2 OpenRouter · z-ai/glm-5.2 · maximum reasoning · Full + SWE + Hard measured 80.50 77.35#17 89.58#1 74.58#18 $2.768 100.0% Result
14 Claude Sonnet 5 OpenRouter · extra-high reasoning · Full + SWE + Hard measured 80.44 76.72#18 78.08#12 86.53#12 $2.812 100.0% Result
15 DeepSeek V4 Pro DeepSeek direct API · maximum-reasoning Hard IQ 79.37 80.06#11 79.68#11 78.38#16 $0.335 100.0% Result
16 GPT‑6 Astra ChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured 78.29 77.86#16 61.92#21 95.08#4 $6.776 100.0% Result
17 Qwen3.7 Max OpenRouter · extra-high reasoning · Hard Intelligence 77.52 73.67#21 88.99#3 69.90#20 $0.906 100.0% Result
18 MiniMax M3 OpenRouter · extra-high reasoning 77.19 77.92#15 86.88#6 66.77#21 $0.182 100.0% Result
19 GPT‑6 Sol ChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured 76.09 83.13#9 52.56#25 92.59#6 $1.310 100.0% Result
20 GPT‑6 Luna ChatGPT Codex subscription · extra-high reasoning · Full + SWE + Hard measured 74.99 78.57#13 62.59#20 83.82#13 $0.070 100.0% Result
21 MiniMax M3 Direct Plus MiniMax direct API · extra-high reasoning · costed at MiniMax's published M3 API rates 70.51 71.96#23 67.77#18 71.80#19 $0.169 100.0% Result
22 Kimi K2.7 Code OpenRouter · Kimi K2.7 Code · extra-high Hard IQ 68.18 79.47#12 58.61#23 66.46#22 $0.732 100.0% Result
23 Step 3.7 Flash OpenRouter · stepfun/step-3.7-flash · extra-high reasoning · Full + SWE + Hard measured 67.16 87.09#3 80.39#10 33.99#27 $0.494 100.0% Result
24 NVIDIA Nemotron 3 Ultra OpenRouter · nvidia/nemotron-3-ultra-550b-a55b · extra-high reasoning 63.45 67.15#26 58.63#22 64.56#23 $0.489 100.0% Result
25 Gemma4‑12B‑Coder Fable5/Composer2.5 Q4_K_M Local model Local GGUF · llama.cpp Vulkan · Q4_K_M 57.62 69.89#24 54.01#24 48.96#25 $0 100.0% Result
26 Qwythos‑9B Claude Mythos Q8_0 Local model Local GGUF · llama.cpp Vulkan · Q8_0 · 256K allocation verified 52.86 68.15#25 46.51#26 43.91#26 $0 100.0% Result
27 Ornith‑1.0‑35B Q4_K_M Local model Local GGUF · llama.cpp Vulkan · Q4_K_M · 35B MoE 50.11 61.58#27 39.16#27 49.58#24 $0 100.0% Result

A tactical testbed, not a latency race.

Resyst Arena evaluates spatial strategy in deterministic turn-based games. Each public match summary links to a replay with board states, legal actions, events, and tactical telemetry.

  • Side-swapped series before ranking-grade claims.
  • Legal action rate and invalid actions are first-class evidence.
  • Replay data stays linked to tactical summaries.
Open the Arena replay room
Figure 3 DeepSeek V4 Flash vs Kimi K2.7 Code Final position after 60 turns. Kimi K2.7 Code won by adjudication. Core integrity 27 for DeepSeek V4 Flash (A) and 29 for Kimi K2.7 Code (B), out of 30.
  • Side A core
  • Side B core
  • Unit, W worker, S striker
  • Resource
  • Control zone

Scores are claims with receipts.

Separate lanes

Agentic, software-engineering, and Hard Intelligence diagnostics are preserved as distinct measurements before any publication formula combines them.

Runtime honesty

The same model can appear through different providers or runtimes. The table exposes basis metadata instead of hiding infrastructure differences.

Evidence thresholds

Single runs are evidence records. Stronger claims require repeated series, side swaps, seed variation, and comparable scoring settings.

Public by design. Auditable by default.

Benchmark summaries are published as versioned data files. The presentation layer is intentionally separate from the scoring harness, so rankings can evolve without rewriting the public record.