Benchmark evidence
OSWorld 2.0
AutoBot with GPT-5.6 Sol Max reached 32.41% binary accuracy and 64.28% partial accuracy across all 108 tasks.

Scoped comparison
On OSWorld 2.0 (release 2026.08.08, all 108 tasks), AutoBot with GPT-5.6 Sol Max reached 32.41% binary accuracy in a self-run evaluation: 18.5% above the official GPT-5.6 Sol Max result (27.34%) and above Claude Opus 5 Max (31.43%).
Method and verifier
The final aggregate selects each task’s best valid attempt across the preserved run batches. The repository publishes the task results, summary, evidence, checksums and a local verifier.
This is a self-run evaluation compared with a published leaderboard snapshot. It is not an official leaderboard submission.