Rank #9
Model result, rank 9 of 27
GLM 5.3 Flash
z.ai Coding Plan · glm-5.3-flash · maximum reasoning with thinking enabled · Full + SWE + Hard measured. Public result card with the model’s overall score, lane measurements, runtime and cost telemetry, and the ranking formula.
- GLM 5.3 Flash
- Cohort median
- Cohort best
Full rank #19
SWE rank #16
Hard rank #3
99.2% reliability
All-around publication view
The overall score averages the measured major lanes while keeping each source measurement visible.
Full / Agentic benchmark
This lane captures instruction following, structured behavior, tool discipline, and general agentic reliability.
Software engineering MVP
This lane is closer to implementation usefulness: source handling, architecture cleanliness, and deliverable quality.
Hard Intelligence diagnostic
Hard Intelligence measures active inquiry, online adaptation, evidence-driven self-repair, and authority/salience integrity.
Runtime economics
Cost, time, and token basis are normalized telemetry. They explain tradeoffs; they do not overwrite the capability score yet.
Why the result lands here.
The result is best read as a balanced benchmark entry: one overall score plus the lane measurements that produced it. Hard Intelligence score is 95.78 and contributes to the overall score alongside Full/Agentic and SWE. Runs on the z.ai Coding Plan subscription rather than a metered API key. Reasoning effort is max with thinking enabled on every lane, which is the strongest setting this route exposes. Costed at z.ai’s published rates ($0.15/1M input, $0.50/1M output, $0.03/1M cached input) as a comparable basis, not an invoice. Implicit prompt caching is billed at the cached rate. Thinking cannot be disabled on this route, so reasoning tokens dominate the output; on the Hard diagnostic they were 93.9% of all emitted output tokens. Hard Intelligence is a public diagnostic, not a hidden official score. Token median and P90 per scored item are intentionally omitted: their derivation could not be reproduced on the same basis as the published rows.