Skip to main content
Back to timeline
arXivSource publication:

Eleven LLMs differ in which information they judge worth pursuing, producing markedly different breadth and depth in open-ended elicitation trajectories

Synopsis

Using 142 arguments from IBM-Rank-30k as a shared information space, the study feeds 11 open-weight instruction-tuned LLMs' judgments of whether a candidate piece of information is worth pursuing (the probability of an affirmative label, treated as an information-seeking score) into one sequential elicitation simulation with a fixed selection rule, fixed respondent behavior, and no natural-language question generation, characterizing how model-specific information-seeking preferences shape the breadth and depth of elicitation trajectories and testing the effects of interaction history, candidate-set size, response labels, and prompt wording.

Source-provided article image: On Open-Ended Information Seeking for Information Elicitation Agents
Figure 1 ·

Figure 1: Illustration of the experimental elicitation setup and example trajectories.

arXiv

Interpretation

Over the same information space and elicitation objectives, models assign substantially different but not arbitrary information-seeking scores: most pairwise correlations fall between .20 and .60, P4 shows consistently lower agreement (.09–.22), and correlations with semantic similarity are low for all models (.14–.34). Prior work largely frames information seeking as choosing questions that reduce uncertainty about an unknown target, or values information by its contribution to a downstream decision; this work measures the model's own judgment of which information is worth pursuing in open-ended elicitation and shows that judgment is not explained by semantic relatedness alone. 11 open-weight instruction-tuned models across five families and multiple parameter scales; shared information space of 142 arguments randomly sampled from IBM-Rank-30k (two per each of 71 topics), each argument evaluated under all 71 topic objectives with no interaction history; evaluations are deterministic forward passes and were not repeated across random seeds.

Plugging these scores into a sequential simulation with a fixed selection rule, changing only the scoring model yields markedly different breadth–depth behavior: topic-switch rates from .295 (P4) to .560 (Q14), distinct-topic counts from 23.4 to 44.7, and mean same-topic run lengths from 1.76 to 3.35; most models also select higher-quality arguments than random. The simulation separates the evaluation of prospective information from question generation and respondent behavior, letting model-specific preferences be studied while other components are held fixed, and shows that local preferences compound into trajectory-level patterns. 100 trajectories of 100 steps per condition across three random seeds, 30,000 elicitation steps in total; candidate-set size k=7, hypothesis confirmation probability fixed at 0.5, responses sampled independently of hypothesis, topic, history, and score; random selection serves as a trajectory baseline.

Increasing candidate-set size k from 3 to 13 expands elicitation breadth (L8 from 29.7 to 40.0 topics, Q14 from 35.4 to 48.2, M8 from 26.7 to 35.2) without removing model ordering differences; the best available argument quality rises with rapidly diminishing returns (about .935 to .984 to .997), selected quality rises more gradually, and the rate of selecting the maximum-quality candidate generally falls with k. Candidate-set size is characterized as an opportunity constraint: it changes both the range and the quality of available information, while model preferences continue to determine which opportunities are pursued. Representative models from three families compared at k=3, 5, 7, 9, 11, and 13; argument quality uses IBM-Rank-30k human-annotated weighted-average scores as an external property, not as the elicitation objective or a preference measure.

When the prompt explicitly states that repeated information is less valuable, scores for candidates highly similar to previously confirmed information drop sharply (Q14 from .823 to .257, M14 from .865 to .634, M24 from .931 to .708) and propagate to selection rates (Q14 from .211 to .042, M14 from .180 to .053, M24 from .167 to .049); with a neutral prompt that does not mention repetition, these reductions largely disappear. Redundancy sensitivity needs no separate deduplication, similarity threshold, or candidate filtering; it operates through the same information-seeking score, and providing interaction history alone does not reliably make models discount repeated information—how the judgment is specified matters. At each step the same candidates are scored with and without topic-specific history; similarity is measured with cross-encoder/stsb-roberta-large and candidates are grouped by high versus low similarity, distinguishing whether overlapping information was previously confirmed or only unconfirmed; Appendix D shows candidate–history pairs with similarity up to .97; the prompt ablation covers Q7, Q14, M8, M14, and M24.

Perspective

The work targets agentic elicitation where elicitation decisions are delegated to a foundation model, and applies to open-ended dialogue that requires continually judging which information is worth pursuing, such as recruiting, journalism, and oral history. It lets researchers compare model-specific information-seeking preferences under a fixed selection rule and observe how those preferences compound into trajectory-level breadth and depth; for practitioners, candidate-set size k is characterized as a tunable opportunity constraint, where increasing k expands both the range and quality of available information with diminishing returns, and redundancy discounting can be achieved by making repetition explicitly relevant to the assessment without a separate deduplication or filtering module. The results apply to controlled settings where candidate–objective evaluation is the interface and respondent behavior is abstracted away.

Information-seeking scores are interpreted as model-specific preferences rather than directly comparable measurements of information value, since affirmative-label probabilities need not be calibrated across models; the respondent confirms with a fixed probability independently of everything else, so real human response behavior is not modeled and no consistency constraints are imposed across confirmed views; the simulation does not generate natural-language questions, so question wording and dialogue dynamics are not covered; under the alternative Valuable/Not valuable labels most models show mean ranking correlations above .70 (five above .88) but G12 is .459, and Appendix C shows label differences can propagate to trajectory-level behavior, with M8 changing substantially; the candidate-set size analysis reports only representative models from three families; and the findings rest on IBM-Rank-30k controversial-topic arguments, so transfer to other information spaces and real interactions remains to be tested.

Sources