Rank #1
Model result, rank 1 of 27
DeepSeek V4.1 Flash
OpenCode Go relay · Full + SWE + Hard measured. Public result card with the model’s overall score, lane measurements, runtime and cost telemetry, and the ranking formula.
- DeepSeek V4.1 Flash
- Cohort median
- Cohort best
Full rank #2
SWE rank #2
Hard rank #10
100.0% reliability
All-around publication view
The overall score averages the measured major lanes while keeping each source measurement visible.
Full / Agentic benchmark
This lane captures instruction following, structured behavior, tool discipline, and general agentic reliability.
Software engineering MVP
This lane is closer to implementation usefulness: source handling, architecture cleanliness, and deliverable quality.
Hard Intelligence diagnostic
Hard Intelligence measures active inquiry, online adaptation, evidence-driven self-repair, and authority/salience integrity.
Runtime economics
Cost, time, and token basis are normalized telemetry. They explain tradeoffs; they do not overwrite the capability score yet.
Why the result lands here.
Rank #1 is the current all-around reference point: strong Full/Agentic performance, competitive SWE delivery, and transparent runtime telemetry. Hard Intelligence score is 88.11 and contributes to the overall score alongside Full/Agentic and SWE. Relay model: deepseek-v4.1-flash on the OpenCode Go relay. No effort level is published: the relay reports effectively flat reasoning-token counts across levels, so no effort label is claimed. Hard Intelligence is a public diagnostic, not a hidden official score.