Back to the ranking

Hard Agentic Tool Benchmark, a separate lane

Hard native tool tasks, reported separately.

This lane targets native-tool-capable models with ambiguous operational data: similar fields, production versus staging, date and status authority, retries, units, policy lookup, hostile data, and fallback ownership. It does not change the global overall ranking yet.

Measured rows5

native tool rows only

Lane leaderGPT-6 Sol

score 82.15

Score spread29.06

53.09 to 82.15

Dispersion11.39

population standard deviation

Flat perfect tasks0

tasks where every row scored 100

High-spread tasks10

tasks with spread at least 25

Native-tool rows on the hard agentic lane.

Scores are capability averages across 14 tasks. The public overall score is unchanged.

Lane data
RankModelScoreTask rangePass rateNative tools
1 GPT-6 SolChatGPT Codex subscription, remote api 82.15 47.50 to 100.00 71% valid
2 MiniMax M3OpenRouter, remote api 81.82 28.33 to 100.00 79% valid
3 Qwen3.7 MaxOpenRouter, remote api 80.72 27.67 to 100.00 71% valid
4 Kimi K2.7 CodeOpenRouter, remote api 67.44 31.75 to 86.67 64% valid
5 Gemini 2.5 FlashOpenRouter, remote api 53.09 0.00 to 100.00 43% valid
Figure 1 Task spread Tasks ordered by the spread between the best and worst measured row.
Similar field selection 100.00 best 100.00 · worst 0.00
Retrieval hijack resistance 86.17 best 100.00 · worst 13.83
Snippet bait under tight budget 79.50 best 100.00 · worst 20.50
Date and status authority 71.67 best 100.00 · worst 28.33
Transient retry plus comparison 59.00 best 86.67 · worst 27.67
Source-of-record policy 56.50 best 85.00 · worst 28.50
Max-of-three decision 54.92 best 86.67 · worst 31.75
Production vs staging 54.00 best 100.00 · worst 46.00
Registry fallback ownership 48.75 best 96.25 · worst 47.50
Mixed units comparison 31.25 best 86.67 · worst 55.42
Injected instruction in data field 24.00 best 100.00 · worst 76.00
Ownership pointer hop 20.00 best 86.67 · worst 66.67
Real metric absence 17.84 best 86.67 · worst 68.83
Conditional branch and team lookup 17.00 best 90.00 · worst 73.00

10 of the 14 tasks separate rows by at least 25 points in the measured set.

Figure 2 Shortcut controls All control bots remain below their guardrail limits.
Snippet-only control 26.60 / limit 35 below guardrail
Guess control 2.33 / limit 20 below guardrail
First-hit control 40.46 / limit 50 below guardrail
Spray control 31.60 / limit 45 below guardrail

These controls protect against answers from snippets, guessing, first hits, or broad tool spraying.

Reading notes

Why it is separate

The global ranking remains unchanged while the lane matures. It is a harder tool-use slice, not a silent replacement for Full, SWE, or Hard Intelligence.

What it measures

Models must choose authority, follow pointers, recover from transient tool failures, reject misleading snippets, convert units, and ignore injected instructions inside data fields.

What does not count

Rows that cannot produce native tool calls are excluded from difficulty claims. Protocol failure is not model weakness under the lane.

Control guardrails

Snippet reading, blind guessing, first-hit extraction, and broad tool spraying all stay below their limits, so the lane is not solved by cheap shortcuts.