Skip to main content
Back to timeline
arXivSource publication:

STAVE replaces in-context demos with two task vectors, beating per-layer interventions on six multimodal models at zero-shot inference cost

Synopsis

The work proposes STAVE, which replaces in-context demonstrations with two task-specific vectors added to the input embeddings of a frozen LLM or LMM, a readout vector updating the answer-producing tokens and a context vector updating the other structural token groups, both trained with answer labels on prompts with and without demos; across six LMMs and five LLMs, STAVE matches or outperforms per-layer interventions with far fewer task parameters, surpasses 15-shot ICL and prior task vectors on 18 text tasks, and runs at zero-shot inference cost.

Source-provided article image: Two Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings
Figure 1 ·

Figure 1: STAVE performance and efficiency. Bars are relative to STAVE, and params use a log scale (Appendix E.3 ).

arXiv

Interpretation

STAVE places the task state at the input embeddings and learns only two vectors: a readout vector covering the final prompt token and the answer cue, and a context vector covering the other five structural token groups, so task parameters do not grow with model depth. Prior demo-free methods place the state at a location searched per task or at every decoder layer, where parameters grow with depth; STAVE needs no layer, head, or token search and does not change prompt length. The paper justifies the design with a first-order loss analysis and a margin bound, and evaluates on six LMMs and five LLMs; in Table 1 STAVE stores 0.008M parameters against 0.13M for LIVE, 0.26M for MimIC, and 2.2M for HiFICL.

On Idefics2-8B and Qwen-VL-7B, STAVE has the best score in every column of Table 1, averaging 82.75% and 83.61%, above MimIC at 81.22% and 82.07% and HiFICL at 80.84% and 82.54%. Those compared methods learn parameters inside decoder layers, whereas STAVE adds two vectors at the input embeddings and still scores higher on average, indicating that fitting the values to the answer matters more than choosing where to add them. Table 1 reports the mean and standard deviation of three runs, with STAVE at 0.19 and 0.14 average standard deviation, lower than most compared methods; Table 2 and Appendix G.1 repeat the comparison on more LMMs.

On the 18 tasks of the TV benchmark with five LLMs, STAVE averages 91.5%, above 15-shot ICL at 90.8% and SITE at 91.0%, while using no demos at inference. Extracted task vectors such as TV and FV average only 75.8% and 48.2%, and LoRA, prefix tuning, and PT train more parameters yet stay below 5-shot ICL. Table 3 covers Pythia 2.8B/6.9B/12B, LLaMA 7B, and GPT-J 6B, weights each task by its number of test queries, and selects configurations on the demo-free development split.

Ablations show the gain comes from the readout/context separation: a single vector, a modality split (Text/image), or moving the answer cue into the context vector (Last/rest) all do worse, and three or seven vectors add no gain. This rules out more vectors or a modality split as the explanation, and matches the prediction of Proposition C.2(iii) that the advantage of two vectors grows as the cosine between the two token sets' gradients falls. Table 14 compares eight group-to-vector configurations on Idefics2 and LLaVA for COCO and VQAv2; Figure 8 measures the pre-training cosine on 18 tasks and five LLMs, where STAVE's advantage is 43.7 points in the lowest third of cosines against 27.5 points in the highest.

Perspective

The method targets researchers and engineering teams who want to adapt a frozen LLM or LMM at zero-shot inference cost: each task needs one training run producing two vectors, added once during demo-free prefill, with prompt length, KV cache shape, and decoder computation equal to zero-shot inference. It applies to prompts with one query image and short answers or captions, and to templated text tasks such as the TV benchmark; training requires access to the model weights and backpropagates through the frozen decoder. The paper names interleaved multi-image and video prompts, and long-form reasoning tasks, as natural next steps.

Several points remain worth watching: training is a one-time cost but backpropagates through the frozen decoder, so memory and time grow with the backbone and weight access is required; the method learns from labeled training samples and loses a few points with a tenth of them. The LMM experiments cover one query image with short answers or captions, leaving multi-image, video, and temporal-understanding tasks untested. The division of labor between the readout and context vectors is characterized on Idefics2 and LLaVA checkpoints, so the mechanistic detail on other backbones remains an open question.

Sources