ThinkingBox: grading agents by backend state across 507 workflows

Microsoft and Hugging Face published ThinkingBox on October 3, 2026, a sandbox and benchmark that evaluates AI agents by the terminal backend state and side effects they produce. The environment runs 507 stateful business workflows, each attempt isolated and repeated 20 times, and uses executable checks to compare actual database changes against required end states.

What ThinkingBox measures

The benchmark treats tool calls and final responses as proxies and grades the records agents leave behind. In a 121,680-trial ablation across 12 LLM models, 79,853 attempts failed executable checks. Of those failures, 67.24% still terminated cleanly and reported no final tool error; executable judges found wrong field values in 77.61% of failed runs, unintended extra effects in 43.30%, and missing required effects in 25.36%.

ThinkingBox reports three metrics per task: pass@1 (share of attempts that succeed), pass@20 (share of tasks solved at least once in 20 tries), and observed 20/20 (tasks that passed all 20 trials). The authors emphasize that a single correct run does not imply reliability: one success is not reliability.

Key findings on models, costs and failures

On pass@1, Claude Opus 5.5 led overall at 67.16%, with Claude Opus 5 at 66.50% and GPT-5.4 at 65.36%. The strongest open-weight model reported was Kimi-K3 (57.37% overall), while domain differences were large (for example, Claude Opus 4.6 scored 68.62% on retail but 8.30% on auto insurance).

Consistency varied: GPT-6 Astra retained 78% of its single-attempt rate under repetition, while Claude Opus 5.5 and Claude Opus 5 each retained about 71%. Kimi-K3 solved the largest number of tasks at least once (476 of 507) but only completed 68 tasks in all 20 attempts.

The authors also assessed cost-efficiency. For single successful attempts, GPT-5.4 was estimated at $0.131 per successful task attempt in a 507-attempt sample. For dependable tasks (passed 20/20), GPT-5.4 cost an estimated $6.80 per dependable task (128 tasks passed 20/20), while GPT-6 Astra and Claude Opus 5.5 cost about $7.45 and $7.80 per dependable task, respectively.

Failure signatures were dominated by tool-handling issues: the ThinkingBox paper reports 79.9% of failures as tool usage problems, 10.3% as wrong state updates, 7.0% as incomplete user resolutions, and 2.9% as no state-changing action.

How it runs and how to try it

Each task in ThinkingBox defines a starting backend state, available MCP tools, domain policy and executable checks. Attempts run in isolated MCP sessions; a side-effect extractor derives actual changes and deterministic judges compare them to required end states. The release includes the sandbox and ThinkingBox-Bench behind the OpenEnv interface.

ThinkingBox is available through Hugging Face and the OpenEnv adapter. The release was tested on Linux and WSL and requires Python 3.11+, uv, Docker, a thinkingbox-data checkout at the pinned release, and model endpoints for the agent, simulated user and judge.


Original source: Hugging Face Blog

Leave a Comment