Benchmark evidence

OSWorld 2.0

AutoBot with GPT-5.6 Sol Max reached 32.41% binary accuracy and 64.28% partial accuracy across all 108 tasks.

OSWorld 2.0 comparison with AutoBot at 32.41% binary accuracy

Scoped comparison

On OSWorld 2.0 (release 2026.08.08, all 108 tasks), AutoBot with GPT-5.6 Sol Max reached 32.41% binary accuracy in a self-run evaluation: 18.5% above the official GPT-5.6 Sol Max result (27.34%) and above Claude Opus 5 Max (31.43%).

Method and verifier

The final aggregate selects each task’s best valid attempt across the preserved run batches. The repository publishes the task results, summary, evidence, checksums and a local verifier.

Inspect the benchmark folder

This is a self-run evaluation compared with a published leaderboard snapshot. It is not an official leaderboard submission.