Rank #3
Model result, rank 3 of 27
GPT‑5.5
ChatGPT Codex subscription · gpt-5.5 · extra-high reasoning · Full + SWE + Hard measured. Public result card with the model’s overall score, lane measurements, runtime and cost telemetry, and the ranking formula.
- GPT‑5.5
- Cohort median
- Cohort best
Full rank #7
SWE rank #8
Hard rank #8
100.0% reliability
All-around publication view
The overall score averages the measured major lanes while keeping each source measurement visible.
Full / Agentic benchmark
This lane captures instruction following, structured behavior, tool discipline, and general agentic reliability.
Software engineering MVP
This lane is closer to implementation usefulness: source handling, architecture cleanliness, and deliverable quality.
Hard Intelligence diagnostic
Hard Intelligence measures active inquiry, online adaptation, evidence-driven self-repair, and authority/salience integrity.
Runtime economics
Cost, time, and token basis are normalized telemetry. They explain tradeoffs; they do not overwrite the capability score yet.
Why the result lands here.
This is a top-tier all-around entrant: the aggregate score remains close to the leader, with lane-level tradeoffs shown separately. Hard Intelligence score is 88.28 and contributes to the overall score alongside Full/Agentic and SWE. Runs on the ChatGPT Codex subscription; the transport is the Codex Responses route rather than a metered API key. Reasoning effort is extra-high on every lane. OpenAI documents none/low/medium/high/xhigh for gpt-5.5, so this is the model's ceiling and no higher level exists to leave unused. Costed at the model’s published Standard API rates ($5.00/1M input, $30.00/1M output) as a comparable basis, not an invoice: the subscription bills no marginal per-token charge. Hard Intelligence is a public diagnostic, not a hidden official score. Token median and P90 per scored item are intentionally omitted: their derivation could not be reproduced on the same basis as the published rows.