Research dataset · Web + Game + Mobile + DevOps CUA frontiers

CUA-SWE
Corpus Viewer

An evidence-first view of web, game-debugging, Mobile, and DevOps tasks that distinguish code-only agents from agents that can inspect and operate running software.

105Total tasks
36Web domain
29Game domain
20Mobile domain
20DevOps domain

01 · Corpus domains

Four domains, one evaluation vocabulary.

Tasks are grouped first by software domain: Web, Game, Mobile, and DevOps. Web reports each measured model and condition separately. Other domains retain their published Easy/Lower and Frontier/Upper classifications.

A

Web domain

36 tasks

All tasks in denominator · Results by model and condition. Browser-based frontend debugging on the frozen evaluated Web36 release.

B

Game domain

29 tasks

11 Easy · 18 Frontier. Play-to-Debug tasks across twelve games with matched code-only and CUA evidence.

C

Mobile domain

20 tasks

11 Lower · 9 Upper. MobileGym-derived tasks with formal GPT-5.6 CUA runs and exact-gold visual replays.

D

DevOps domain

20 tasks

15 Lower · 5 Differential / Upper. Operator-console debugging with deterministic protected-runtime verification.

02 · Evidence explorer

GUI evidence and coding activity, side by side.

Screenshot-derived videos show only computer-use interaction. Terminal commands, source edits, tests, and agent reasoning appear in a separate ordered activity stream sourced from either the exact historical run or an explicitly labeled replacement rerun.

Select a task

Inspect its paired evidence.

03 · Performance and curation

The domain-first corpus ledger.

Compare Web, Game, Mobile, and DevOps tasks, then refine by difficulty and collection.

Showing the curated corpus
Compare Task Domain Difficulty Collection Code-only CUA Source Evidence

Task comparison

Evidence side by side