Skip to main content
Back to timeline
arXivSource publication:

How a Small GUI Grounding Model Should Receive the Action Type: Auxiliary Loss, Additive Embedding, and Prompt Word Each Gain 5–7 hit@0.10 Points, but Much of the Gain Comes from a Serialization Artifact

Synopsis

Fine-tuning Qwen2-VL-2B with LoRA on Android in the Wild, the study compares five ways of supplying the action type to a small GUI grounding model under matched data, compute, and decoding, finding that an auxiliary loss, an additive embedding, and a prompt word each raise hit@0.10 by five to seven points, but after retraining without the degenerate class created by the serializer no mechanism beats the baseline on hit rate, and only the auxiliary loss shortens the average miss.

Source-provided article image: Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type?
Figure 2 ·

Figure 2: Conditioning advantage over the flat baseline as a function of training-set size on all_with_coords , three seeds per cell, 95% episode-clustered bootstrap intervals. Each size is scored on a different validation slice with a different class mix, so only within-size differences are meaningful, and 24 contrasts are shown without correction.

arXiv

Interpretation

On a stream mixing clicks, scrolls, and type events, the auxiliary loss, additive embedding, and prompt word each beat the flat baseline by five to seven hit@0.10 points and are indistinguishable from one another; hard routing and the prepended token are within noise of the baseline. Prior work lacked a comparison of type-delivery mechanisms under matched data, compute, adapters, and decoding, and lacked a test of whether the trained model reads the signal at all. Five seeds, episode-clustered bootstrap intervals, and seed-level paired tests; seed variance is large (baseline spans 0.169 to 0.280), motivating paired statistics rather than a best run.

Much of the mixed-stream gain comes from the authors' own serialization: AITW records type events with an off-screen touch point that the serializer clamps to the origin, so the flat model must predict a fixed point for a class it is not told about, and that class pulls the baseline's clicks toward the origin more broadly. Attributes an apparent spatial-prior gain to a preprocessing-induced degenerate class and quantifies it by retraining without that class. Retraining without type events raises the flat baseline from 0.229 to 0.296 hit@0.10 and leaves every conditioned variant within 0.025 with intervals through zero; on the clean stream no mechanism beats the baseline on the headline metric, with intervals tight enough to exclude a two-point effect.

Interventions on trained models show both learned embeddings are read at inference: a wrong type collapses the additive model from 0.275 to 0.054 and the prepended token from 0.242 to 0.079; zeroed, the additive model stays within a point of the baseline while the prepended-token model falls seven points below it. Introduces a wrong, zero, and class-mean intervention protocol that separates training effects from inference effects and tests whether a conditioning signal is read. Five seeds with episode bootstrap and seed-level tests; the prepended token's rows barely move at the shared learning rate (displacement 0.016 from an initialization of norm 0.78) and reach the level of the other three mechanisms when trained ten times faster.

Injecting conditioning through inputs_embeds makes Qwen2-VL fall back to one-dimensional positions for its image tokens, costing nine points, a silent failure that had inverted the authors' conclusion. Documents and traces a silent positional-encoding fallback against the library source, showing how it reversed a conclusion. Inferred from the library source and a frozen-embedding control rather than from position ids in the failing runs; a forward hook with input_ids intact removed the deficit.

Perspective

The results are aimed at researchers and engineers fine-tuning small VLM grounding models on AITW-like data, in settings where the action type and coordinate are emitted in a single autoregressive stream. Their value lies in a reusable comparison and intervention protocol: compare type-delivery mechanisms under matched data, compute, adapters, and decoding, and use wrong, zero, and class-mean interventions to test whether the signal is read. For deployment, the prompt word and auxiliary loss achieve effects comparable to learned embeddings without new parameters, and the additive embedding stays within a point of the baseline when zeroed, making it a low-risk default. The serialization-audit lesson transfers directly to any pipeline that supervises a coordinate.

Whether the residual effect of the auxiliary loss in shortening the average miss on the clean stream holds on larger slices, and whether the degenerate-class hazard is a property of one serialization and one dataset, remain open. The prepended token's failure at the shared learning rate is attributed to the schedule rather than the architecture, but the faster-trained token uses three seeds and no frozen-slot control was trained on the fixed path. The class-mean intervention is matched to the prepended token in direction but not in norm. In the end-to-end result, the predicted-type pipeline's margin over the baseline is not established, and a wrong type collapses every model conditioned at inference, making the classifier the main risk. The attention analysis uses one seed and one layer, and teacher forcing makes the measurement not directly comparable to grounding decided during prefill.

Sources