Rank #24
Model result, rank 24 of 27
NVIDIA Nemotron 3 Ultra
OpenRouter · nvidia/nemotron-3-ultra-550b-a55b · extra-high reasoning. Public result card with the model’s overall score, lane measurements, runtime and cost telemetry, and the ranking formula.
- NVIDIA Nemotron 3 Ultra
- Cohort median
- Cohort best
Full rank #26
SWE rank #22
Hard rank #23
100.0% reliability
All-around publication view
The overall score averages the measured major lanes while keeping each source measurement visible.
Full / Agentic benchmark
This lane captures instruction following, structured behavior, tool discipline, and general agentic reliability.
Software engineering MVP
This lane is closer to implementation usefulness: source handling, architecture cleanliness, and deliverable quality.
Hard Intelligence diagnostic
Hard Intelligence measures active inquiry, online adaptation, evidence-driven self-repair, and authority/salience integrity.
Runtime economics
Cost, time, and token basis are normalized telemetry. They explain tradeoffs; they do not overwrite the capability score yet.
Why the result lands here.
The model is stronger in the Full/Agentic lane than in the SWE lane; the overall score is therefore shown with both component lanes visible. Hard Intelligence score is 64.56 and contributes to the overall score alongside Full/Agentic and SWE. Hard Intelligence is published alongside Full/Agentic and SWE for the current ranking. Full long-context probe rescored under suite v4 (rubric: correct answer zeroed by the phrase list).