About ACU Gym
A public experimental environment for practicing and evaluating computer-use agents on interactive web missions.
What “ACU” means here
ACU GYM is a synthetic training and evaluation environment for Agentic Computer Use. The ACU name originated internally as Arkhē Computer Use and is used publicly here as Agentic Computer Use. ACU is this project’s descriptive name for the category, not an established industry-standard acronym.
Use Computer Use agent for an individual agent. Use Agentic Computer Use for the broader capability and evaluation domain.
Claim boundary
ACU GYM is a functional synthetic curriculum and experimental evaluation harness for computer-use agents. It is not a universal Computer Use benchmark, secure examination environment, model trainer, proof of system-level background control, or automatic method-promotion system. There is no independent external attestation by default. Results apply only to the missions and execution modes shown.
This mission simulates owned and human workspaces inside one application. It tests target selection, interference avoidance, collision handling and verification logic. By itself it does not prove real background-tab operation, operating-system focus preservation, unchanged physical pointer state or coexistence with a human using the device.
Methodology (honest)
- Guided Practice (engine mode TRAINING) may expose control IDs, coaching, and predicate hints. Results are
GUIDED_TRAINING_RESULT. - Blinded Evaluation exposes only a public task brief (objective, scope, budgets, completion instruction). No solution selectors.
- Execution mode must be declared (HEADED_UI, SEMANTIC_DOM, API_DIRECT, …). UI-only exams reject API transport.
- Receipts use SHA-256 digests for deterministic integrity checking — not signatures or secure audit chains.
- Shared Desk Simulation is within-page. External attestation is required for
EXTERNALLY_ATTESTED_SHARED_DESK. - Held-out seeds use commitment/reveal. Client-side only — not secret from a machine owner.
- Scoring uses
scoringVersion=2for new runs. Historical receipts keep original v1 composites immutably. v1 and v2 are not directly comparable.predicateVerificationis same-engine;independentVerificationis not observed unless a distinct external verifier is present. - Contaminated demo seeds (e.g. 9001, leaked evaluation IDs) are permanently rejected as held-out evidence. Details live in the Evaluator console.
Operator tools: Evaluator console.
Mission distinctness snapshot
Missions: 21 · Distinct state-machine topologies: 21 · Distinct primitives: 18
No exact topology duplicates detected in the current catalog.