Researchers extract a user role vector from model activations, finding assistant bias is structurally identifiable and user behavior can be directionally steered
Synopsis
This work analyzes "assistant bias" in LLM user simulators at the level of model activations, extracting a user role vector by contrasting how the model represents user versus assistant perspectives on the same dialogue, and finds that the user direction is identifiable in activations, elicits user-like behaviors, and captures characteristics distinct from assistant traits, while user-role activation associates with simulation realism and steering strengthens it but can also exaggerate user behaviors and override individual user profiles.
Figure 1: We extract the user role vector to analyze (A) how assistant bias shapes LLM simulator behavior and (B) whether steering along this direction influences it.
arXivInterpretation
The user direction is identifiable in model activations, elicits user-like behaviors, and captures characteristics distinct from assistant traits. Prior work outlines that assistant bias is baked in during model training and that role-playing prompts fail to override it; this work moves from the prompt level to the representation level by contrasting user versus assistant perspectives on the same dialogue to extract a user role vector. Evidence comes from contrastive analysis of model activations; the authors report that the user direction is identifiable, elicits user-like behaviors, and is distinct from assistant traits, without reporting specific sample sizes or statistics in the abstract.
User-role activation associates with simulation realism and steering strengthens it, but steering can also exaggerate user behaviors and override individual user profiles. Beyond locating the user direction, the work examines the effect of steering along it, highlighting a tension between enhancing realism and preserving individual user profiles, which is a representation-level characterization of simulator controllability. Evidence comes from observations linking user-role activation to simulation realism and from steering experiments; the abstract does not report specific effect sizes or evaluation metric values.
Assistant bias is structurally identifiable and user behavior can be directionally analyzed. The work provides a representation-level analysis of LLM user simulators, turning assistant bias from a behavioral phenomenon into a structure that can be located and manipulated in activation space. The conclusion rests on the two findings above and is a summary of the activation-level analysis; the abstract provides no cross-model or cross-dataset validation details.
Perspective
This work targets research and engineering settings that use LLM user simulators for large-scale agent evaluation, and is relevant to readers who need to judge whether a simulator faithfully reproduces real user behavior, including frustration and disengagement. It locates assistant bias as a structure identifiable in activation space and shows that steering along the user direction can strengthen simulation realism, providing a starting point for diagnosing and adjusting user simulators at the representation level.
The abstract gives no sample sizes, effect sizes, or specific evaluation metrics, and does not state the range of models or dialogue data used, so how stable the trade-off between steering-enhanced realism and exaggerated user behaviors is remains an open question for readers. Whether the user role vector is similarly identifiable across different models or tasks is also not addressed in the abstract.
