native tool rows only
Hard Agentic Tool Benchmark, a separate lane
Hard native tool tasks, reported separately.
This lane targets native-tool-capable models with ambiguous operational data: similar fields, production versus staging, date and status authority, retries, units, policy lookup, hostile data, and fallback ownership. It does not change the global overall ranking yet.
score 82.15
53.09 to 82.15
population standard deviation
tasks where every row scored 100
tasks with spread at least 25
Native-tool rows on the hard agentic lane.
Scores are capability averages across 14 tasks. The public overall score is unchanged.
| Rank | Model | Score | Task range | Pass rate | Native tools |
|---|---|---|---|---|---|
| 1 | GPT-6 SolChatGPT Codex subscription, remote api | 82.15 | 47.50 to 100.00 | 71% | valid |
| 2 | MiniMax M3OpenRouter, remote api | 81.82 | 28.33 to 100.00 | 79% | valid |
| 3 | Qwen3.7 MaxOpenRouter, remote api | 80.72 | 27.67 to 100.00 | 71% | valid |
| 4 | Kimi K2.7 CodeOpenRouter, remote api | 67.44 | 31.75 to 86.67 | 64% | valid |
| 5 | Gemini 2.5 FlashOpenRouter, remote api | 53.09 | 0.00 to 100.00 | 43% | valid |
10 of the 14 tasks separate rows by at least 25 points in the measured set.
These controls protect against answers from snippets, guessing, first hits, or broad tool spraying.
Reading notes
Why it is separate
The global ranking remains unchanged while the lane matures. It is a harder tool-use slice, not a silent replacement for Full, SWE, or Hard Intelligence.
What it measures
Models must choose authority, follow pointers, recover from transient tool failures, reject misleading snippets, convert units, and ignore injected instructions inside data fields.
What does not count
Rows that cannot produce native tool calls are excluded from difficulty claims. Protocol failure is not model weakness under the lane.
Control guardrails
Snippet reading, blind guessing, first-hit extraction, and broad tool spraying all stay below their limits, so the lane is not solved by cheap shortcuts.