Public articles linked to the same research event.
arXiv In a multi-agent hypothesis-generation workflow, the authors held opening hypotheses fixed and replayed downstream stages, using an LLM-as-a-judge to decide whether paired hypotheses describe the same mechanism: moving from 0 to 1 self-critique round produced 34.5 percentage points of extra mechanism-level divergence while 1 to 5 rounds added 1.3 pp; under controlled perturbations all three equivalence rules were invariant to meaning-preserving edits, but when the initiating event was replaced the LLM-as-a-judge read 83% of pairs as different mechanisms versus 0% for TF-IDF and 8% for embeddings, and changing only the rubric moved this rate from 38% to 96%.
In a multi-agent hypothesis-generation workflow, the authors held opening hypotheses fixed and replayed downstream stages, using an LLM-as-a-judge to decide whether paired hypotheses describe the same mechanism: moving from 0 to 1 self-critique round produced 34.5 percentage points of extra mechanism-level divergence while 1 to 5 rounds added 1.3 pp; under controlled perturbations all three equivalence rules were invariant to meaning-preserving edits, but when the initiating event was replaced the LLM-as-a-judge read 83% of pairs as different mechanisms versus 0% for TF-IDF and 8% for embeddings, and changing only the rubric moved this rate from 38% to 96%.
In a multi-agent hypothesis-generation workflow, the authors held opening hypotheses fixed and replayed downstream stages, using an LLM-as-a-judge to decide whether paired hypotheses describe the same mechanism: moving from 0 to 1 self-critique round produced 34.5 percentage points of extra mechanism-level divergence while 1 to 5 rounds added 1.3 pp; under controlled perturbations all three equivalence rules were invariant to meaning-preserving edits, but when the initiating event was replaced the LLM-as-a-judge read 83% of pairs as different mechanisms versus 0% for TF-IDF and 8% for embeddings, and changing only the rubric moved this rate from 38% to 96%.
In a multi-agent hypothesis-generation workflow, the authors held opening hypotheses fixed and replayed downstream stages, using an LLM-as-a-judge to decide whether paired hypotheses describe the same mechanism: moving from 0 to 1 self-critique round produced 34.5 percentage points of extra mechanism-level divergence while 1 to 5 rounds added 1.3 pp; under controlled perturbations all three equivalence rules were invariant to meaning-preserving edits, but when the initiating event was replaced the LLM-as-a-judge read 83% of pairs as different mechanisms versus 0% for TF-IDF and 8% for embeddings, and changing only the rubric moved this rate from 38% to 96%.