Skip to main content
Back to timeline
arXivSource publication:

Across three domains, 12 LLM agents consistently favor certain sources, and that preference can outweigh requirement satisfaction; hiding source or supplying missing information reduces it

Related research and updates

Synopsis

Studying end-to-end search with 12 agent models across three domains, the work finds each model prefers some sources and avoids others with broad agreement on which; an item satisfying one requirement fewer is selected about two-thirds of the time when it comes from a preferred source while the better item comes from a dispreferred source, but almost never in the reverse case; hiding source information weakens the preference and relabeling an item with a preferred source raises its selection rate; training that rewards better items can make a source a shortcut for requirement satisfaction, and missing information can trigger preconceptions about the source, while supplying missing information or a prompt countering those preconceptions reduces source preference.

Source-provided article image: Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It
Figure 1 ·

Figure 1: Overview. Top: agent setting (§ 2 ). Bottom: measuring source preference (§ 3 ).

arXiv

Interpretation

In end-to-end search across three domains, all 12 agent models show preference for some sources and avoidance of others, and the models largely agree on which sources are preferred. Prior examinations of source preference were largely not in end-to-end search settings and did not systematically compare multiple models across multiple domains; this work places source preference in end-to-end search and compares items from different sources at the same position satisfying the same requirements. Based on comparisons across 12 agent models and three domains, comparing items from different sources that satisfy the same requirements at the same position.

Source preference can outweigh how well items satisfy the request: an item satisfying one requirement fewer is selected about two-thirds of the time when it comes from a preferred source, while the better item from a dispreferred source is almost never selected. This provides direct evidence of a trade-off between preference and requirement satisfaction, indicating selection is not determined by requirement satisfaction alone. Quantified via selection rates in the contrast between an item satisfying one requirement fewer and a better item, reporting about two-thirds versus almost never.

The information identifying an item's source affects selection by itself: hiding it weakens the preference, and relabeling an item with a preferred source raises its selection rate. Shows source preference does not come only from item content; the source label alone can independently change selection behavior. Observed through two interventions, hiding source information and relabeling the source, tracking changes in selection rate.

The work tests two routes to this preference: training that rewards better items can make a source a shortcut for requirement satisfaction, and missing information can trigger preconceptions about the source; supplying missing information or a prompt countering these preconceptions reduces source preference. Attributes source preference to two mechanisms, a training shortcut and preconceptions triggered by missing information, and offers information supply and prompting as interventions that reduce it. Mechanism tests use training reward settings and missing-information conditions, with interventions verified via supplying missing information and a prompt countering preconceptions.

Perspective

The work targets settings where LLM agents conduct end-to-end search on users' behalf, covering three domains and 12 agent models, applicable to tasks such as buying products, booking hotels, or citing papers where agents choose sources for users. Its conclusions are that source preference is pervasive and largely consistent across models, can outweigh requirement satisfaction, and can be weakened by hiding source, supplying missing information, or using a prompt that counters preconceptions. For readers building agent search systems, this suggests treating source labels and information completeness as variables that affect selection outcomes during evaluation and deployment.

The text is at the abstract level and does not give the specific sources per domain, the names of the 12 models, sample sizes, or statistical test details, so how much preference strength varies across domains and models remains to be confirmed in the full text. The relative contribution of the training-shortcut and missing-information-preconception routes, the robustness of information-supply and prompting interventions under different conditions, and the exact procedures of the hiding-source and relabeling experiments are open questions that require reading the full text.

Sources