Skip to main content
Back to timeline
arXivSource publication:

In multi-agent hypothesis generation, the first self-critique round adds 34.5 pp of mechanism divergence while the equivalence rule itself shapes measured diversity

Related research and updates

Synopsis

In a multi-agent hypothesis-generation workflow, the authors held opening hypotheses fixed and replayed downstream stages, using an LLM-as-a-judge to decide whether paired hypotheses describe the same mechanism: moving from 0 to 1 self-critique round produced 34.5 percentage points of extra mechanism-level divergence while 1 to 5 rounds added 1.3 pp; under controlled perturbations all three equivalence rules were invariant to meaning-preserving edits, but when the initiating event was replaced the LLM-as-a-judge read 83% of pairs as different mechanisms versus 0% for TF-IDF and 8% for embeddings, and changing only the rubric moved this rate from 38% to 96%.

Source-provided article image: Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis Generation
Figure 1 ·

Figure 1: Multi-agent workflow for mechanistic hypothesis generation used in this study. An unexpected experimental finding initiates evidence retrieval. Candidate hypotheses are generated in parallel and iteratively refined through self-critique loops between proposer and critic agents. The resulting hypotheses are then consolidated into a final report.

arXiv

Interpretation

The main effect of critique depth appears when critique is first introduced: relative to matched same-depth reruns, 0 to 1 round produces 34.5 percentage points of extra mechanism divergence, 1 to 5 rounds adds only 1.3 pp, and 0 to 5 rounds totals 39.8 pp. Prior evaluations of iterative refinement rely on gold answers or task metrics unavailable in open-ended hypothesis generation; this work uses same-depth reruns as a variability baseline to separate critique-depth effects from run-to-run stochasticity. Four proprietary instances, five repetitions per critique depth, 60 runs, 420 run-pair comparisons, and 4,863 matched hypothesis pairs; the pattern holds across three judge models, with 0-to-1 extra divergence of 34.5 to 37.3 pp.

All three equivalence rules treat meaning-preserving edits (wording rewrites, or chain rewrites using audited biological aliases) as the same mechanism, but diverge when the initiating event is replaced: the LLM-as-a-judge reads 83% of pairs as different, versus 8% for embeddings and 0% for TF-IDF. The work audits comparison methods with perturbations whose intended relation is known, rather than only reporting agreement among methods, revealing that global similarity measures can be less sensitive when a single causally important component changes. Perturbation variants were retained only after deterministic structural validation; at selected thresholds the three methods achieve 85% to 89% raw pairwise agreement, and no evaluated threshold produces complete agreement among them.

LLM-as-a-judge decisions depend on prompt specification: changing only the rubric that defines same mechanism moves the different-mechanism rate for initiating-event swaps from 38% under a no-rubric prompt to 83% under the full rubric and 96% under a step-by-step rubric; the step-by-step rubric also raises sensitivity to middle-step swaps from 13% to 21%. This turns the evaluation criterion from an implicit setting into an explicit experimental variable, showing that the same model has different sensitivity to causal differences under different criteria. Three prompt variants were compared on the same controlled perturbation pairs with judge model, temperature, sampling, and rendered inputs held fixed; exact repetition changed no decisions, while reversing presentation order changed 5.6% of controlled pairs.

Pairwise equivalence judgment is a shared measurement basis for two evaluation questions: it determines both the estimated effect of self-critique and the diversity count of a generated hypothesis set. Prior systems apply different comparison rules to preference ranking and deduplication; this work shows both rest on the same same/different decision, so the measurement choice affects both kinds of conclusions. Formal definitions express divergence as the fraction of matched pairs judged different and diversity as the number of clusters under an equivalence rule, compared across three rules on the same hypotheses.

Perspective

The results apply to evaluating multi-agent mechanistic hypothesis generation without experimental ground truth: given an unexpected experimental finding, the workflow generates causal chains from an initiating event through intermediate biological steps to that finding, revised through self-critique before consolidation. They address system developers and evaluators who must report refinement effects or output diversity, under a setting where opening hypotheses are fixed, only critique depth is varied, and downstream stages are replayed. The controlled-perturbation audit can be used to compare equivalence rules and rubrics and to inform how measurements are chosen and reported.

No biological ground truth is available for the generated hypotheses, so the results characterize how workflow outputs and evaluation methods behave, not whether critique improves scientific correctness or whether a grouping corresponds to the biologically correct mechanism. Measured diversity remains sensitive to evaluator design: TF-IDF and embedding results depend on similarity thresholds, and LLM same/different judgments vary with judge model, prompt specification, and presentation order; because judges and generators belong to the same broad class of LLMs, agreement across judge families does not rule out shared biases such as self-preference or sensitivity to superficial properties. The controlled perturbations probe only a limited set of semantic and causal changes, and broader, systematically constructed perturbation suites with independently specified rubrics would help show which causal distinctions an evaluator captures and how robustly. Post-hoc evaluation also relies on many LLM pairwise comparisons, whose cost grows with the hypothesis set.

Sources