Mobile agents look intelligent when the assignment is tap this button and book that table. Add state, nested navigation, a few dependencies, and the magic trick starts coughing.

The new GMA benchmark evaluates eight frontier models on 300 tasks across seven open-source Android applications. The researchers organize tasks into four difficulty tiers, creating a controlled way to measure what happens as an agent must preserve more context, navigate longer workflows, and reason about application state.

Performance declines substantially as complexity rises. That finding sounds obvious, but the benchmark makes the failure measurable instead of anecdotal. An agent that completes short, visually direct tasks can still be nowhere near reliable enough for work involving multiple screens, delayed consequences, or a decision that depends on an earlier action.

The researchers also ran controlled harness experiments. Context retention and explicit state tracking can improve performance, but the gains vary by model. In plain English, more memory helps, but memory is not a universal patch for weak planning, bad perception, or an inability to recognize that the interface has entered a different state.

GMA is a preprint and has not completed peer review. Its useful contribution is a tougher definition of progress. Stop grading mobile agents on whether they can survive a choreographed demo. Grade them on whether they can finish complicated work, detect when the world disagrees with the plan, and recover without quietly turning one wrong tap into six more.

LaunchPad positionA clean demo proves that an agent can click. Reliable work requires persistent state, context retention, recovery, and enough judgment to survive a task whose next step depends on what already happened.
Reporting standard

This report draws on the linked primary sources and reputable reporting. Company statements are treated as claims until independently demonstrated.