Skip to content
PUBLIC PREVIEWAgentic Computer Use

ACU GymPractice and evaluate computer-use agents in interactive web missions.

Train agents across forms, dynamic interfaces, shared workspaces, recovery tasks and adversarial UI traps. Run guided practice, structured gauntlets and reproducible evaluations with evidence-rich results.

21 missions · deterministic runs · failure-aware receipts

ACU stands for Agentic Computer Use—this project’s descriptive name for interactive agent evaluation, not an industry-standard acronym.

How it works

1. Choose a mission

Select guided practice, a campaign path or a structured gauntlet.

2. Let your agent operate

The agent interacts with dynamic interfaces, state changes and intentional traps.

3. Review the evidence

Inspect outcomes, action economy, retries, collateral effects and verification results.

Built for

Agent developers testing interaction loops
Researchers studying reliability and recovery
Builders comparing strategies and tool paths
Curious users who want to watch an agent work through real interfaces

21

Missions

6

Learning paths

9

Gauntlets

Yes

Evidence-aware runs

Explore

Advanced evaluation

EXPERIMENTAL

For researchers and operators who need blinded runs, commitments and stricter evidence boundaries. Not required for guided practice.

Open Evaluator

Experimental evaluation environment. Results describe performance on these missions only—not universal computer-use ability, security certification or model training.

Full scientific and security disclosures: About · Evaluator / Methodology.