MintAct: A Unified Visual Agent for Digital Environments
Synopsis
MintAct introduces a family of vision-language models at 2B, 4B, and 8B scales that unify UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use in a single set of weights through a shared screenshot-and-pixel observation space, prompt-conditioned per-domain action sets, and balanced cross-domain mixing, backed by asynchronous RL infrastructure hosting hundreds of concurrent instances, reaching 48.9 on OSWorld-Verified, 39.1 on Online-Mind2Web, and 67.0 on AndroidWorld while matching or exceeding size-matched per-domain specialists.
Interpretation
A single model can cover UI grounding, multi-step navigation on mobile, desktop, and web, and visual tool use simultaneously without sacrificing per-domain quality. These capabilities have typically been built, trained, and benchmarked as separate specialists; MintAct covers all of them with one set of weights and validates this at 2B, 4B, and 8B scales. The paper reports comparisons against size-matched specialist baselines across MMBench-GUI, ScreenSpot-v2, UI-Vision, OSWorld-G, AndroidControl, AndroidWorld, OSWorld-Verified, Weblica, Online-Mind2Web, and MM-ToolSandBox, with MintAct-8B attaining the reported best scores on OSWorld-Verified (48.9), Weblica (74.7), UI-Vision (56.6), and OSWorld-G (64.5).
Unification rests on three design choices: a shared observation and grounding space, prompt-conditioned action sets, and explicitly balanced cross-domain mixing. The paper notes that naively merging per-domain action sets or mixing their data lets domains interfere and erodes per-domain quality, so it instead grounds actions by normalized pixel coordinates on raw screenshots and exposes each domain's actions and tools through a domain-specific system prompt. The paper provides a side-by-side table of per-domain action spaces and the system prompts, and states that the training distribution is kept balanced across domains in both supervised fine-tuning and reinforcement learning; ablations show the joint SFT model matches or exceeds per-domain SFT specialists on mobile, desktop, and visual tool use and stays in a similar range on grounding and web.
The multi-stage recipe (high-resolution single-step SFT, low-resolution multi-step SFT, per-domain RL specialists distilled back via rejection sampling, and joint agentic RL) contributes measurable gains at every stage. Ablations separate the stages: the first is valuable for grounding, the second is essential for navigation and tool use, the two are complementary, and RFT distillation lifts navigation and tool use while keeping a single set of weights. The paper reports that Stage 1 alone reaches 59.3 on UI-Vision but only 9.4 on OSWorld-Verified and 15.6 on Weblica; Stage 2 alone recovers OSWorld-Verified to 41.5 and Weblica to 59.1 but drops UI-Vision to 25.0; combining both yields UI-Vision 56.9 and OSWorld-Verified 42.6.
An asynchronous RL framework tailored to unified visual agents trains stably on heterogeneous environment backends while keeping explicit control over the cross-domain training distribution. To address slow GUI interaction, long multimodal trajectories, and mixture drift in mixed-domain training, the paper designs per-domain quotas, quota-normalized backpressure, a dual-clipped surrogate objective, and truncated importance weights. The framework is described as built on verl and rLLM, hosting more than 200 concurrent desktop instances and more than 100 mobile instances during RL training, with a message queue capacity of 448 and a parameter synchronization interval of 5 steps, and is reported to remain stable under noisy environment feedback and off-policy drift.
Perspective
This work targets digital-assistant settings where one model must handle UI grounding, mobile, desktop, and web navigation, and visual tool use at once, especially at compact scales where serving and deployment cost matter. The paper states that joint RL is currently run only on mobile and desktop, and that extending it to web and visual tool use requires sustaining many heterogeneous environments concurrently, which it leaves to future work. Synthetic environments help most where real interaction data is scarce, such as iOSWorld, while real in-domain RL yields larger gains when in-domain real data is available. The paper also notes that visual tool use is currently treated largely as a separate capability, with automatic switching between pixel-level navigation and dynamic tool calling still open.
Joint RL covers only mobile and desktop, so the effect of joint training with web and visual tool use remains an open question; context grows rapidly as screenshots are appended during long-horizon interaction, and more sophisticated context management is not yet included; trajectory segmentation under a dynamic tool registry provides only coarse credit assignment, which the paper describes as an approximation of end-to-end interaction. In addition, this reading is of the full paper text, and some table values appear as placeholders in the text, so exact per-benchmark numbers should be checked against the original tables.
