91.4% on AndroidWorld.
91.4% success across 116 AndroidWorld tasks. Open source, GPT-5, direct accessibility API. Reproducible from the eval harness in our public repo.
Measured, not implied.
106 of 116 tasks completed in the published mobilerun evaluation run.
Tasks completed.
Calendar, Files, Notes, Markor, Audio, Browser, System, Expense, Recipe, Clock, and more.
Manager-Executor.
Dynamic feedback loops between planning and execution. Robust under long-horizon failure modes.
How we compare.
A dated comparison of publicly reported AndroidWorld results. The current mobilerun run uses the accessibility API with vision fallback.
Scores are pulled from their public reports and the AndroidWorld leaderboard. The eval harness is in our public repo so anyone can reproduce. Updated 29.10.2025
Why the a11y tree wins.
Smaller payloads, richer semantics, faster steps.
| Per-step input | Screenshot | Accessibility tree | Difference |
|---|---|---|---|
| Payload size | ~1,024 KB | ~2 KB | 500× smaller |
| Semantic structure | Pixels only | Labels & roles | Higher accuracy |
| Step latency | High | Low | Faster agents |
| Vision fallback | n/a | When tree is incomplete | Robust |
Where we succeeded.
Successful categories from the 106 / 116 run.
Calendar
Event creation, navigation, recurring schedules.
System
Settings, intents, system-level workflows.
Markor
Note editing, search, formatting.
Recipe
Multi-step content workflows with state.
Expense
Form completion, calculation, lookups.
Audio, Browser, Clock, Files, Notes & more
Across the remaining task families.
All 115 tasks.
Open any task to replay its full execution trajectory.
Run it yourself.
The eval harness is open source. Reproduce the result, or build your own benchmark on the same primitives.