Web domain
36 tasksAll tasks in denominator · Results by model and condition. Browser-based frontend debugging on the frozen evaluated Web36 release.
Research dataset · Web + Game + Mobile + DevOps CUA frontiers
An evidence-first view of web, game-debugging, Mobile, and DevOps tasks that distinguish code-only agents from agents that can inspect and operate running software.
01 · Corpus domains
Tasks are grouped first by software domain: Web, Game, Mobile, and DevOps. Web reports each measured model and condition separately. Other domains retain their published Easy/Lower and Frontier/Upper classifications.
All tasks in denominator · Results by model and condition. Browser-based frontend debugging on the frozen evaluated Web36 release.
11 Easy · 18 Frontier. Play-to-Debug tasks across twelve games with matched code-only and CUA evidence.
11 Lower · 9 Upper. MobileGym-derived tasks with formal GPT-5.6 CUA runs and exact-gold visual replays.
15 Lower · 5 Differential / Upper. Operator-console debugging with deterministic protected-runtime verification.
02 · Evidence explorer
Screenshot-derived videos show only computer-use interaction. Terminal commands, source edits, tests, and agent reasoning appear in a separate ordered activity stream sourced from either the exact historical run or an explicitly labeled replacement rerun.
Select a task
03 · Performance and curation
Compare Web, Game, Mobile, and DevOps tasks, then refine by difficulty and collection.
| Compare | Task | Domain | Difficulty | Collection | Code-only | CUA | Source | Evidence |
|---|