Skip to main content
Back to timeline
arXivSource publication:

InvestigationWorlds builds a legal investigation environment from 100 real federal rulings, and the best model solves only one-third of cases

Related research and updates

Synopsis

The authors introduce InvestigationWorlds: they retrieve 100 real U.S. federal civil cases with decided summary judgment motions from PACER and use an attorney-validated generation pipeline to synthesize role-tagged documents around the original record, so that one corpus admits multiple coherent but incompatible factual readings of which only one matches the court-adopted hypothesis; evaluating six frontier agents on 100 cases, they find agents often commit to incorrect hypotheses despite retrieving relevant evidence, with the highest case-solved rate only 33.3%.

Source-provided article image: InvestigationWorlds: An Agentic Environment for Legal Investigation
Figure 1 ·

Figure 1 : The InvestigationWorlds pipeline. Real federal cases with summary judgment motions are filtered from PACER. Statements of facts, supporting exhibits, and the Court’s opinion are processed by an LLM pipeline that extracts entities, ground facts, and the two party hypotheses ( h p h_{p} , h d h_{d} ), and assembles them into a document graph and hypothesis Directed Acyclic Graph (DAG). A planner-and-judge agent then augments the document graph with synthetic exhibits tagged by evidentiary role (see Table 1 ).

arXiv

Interpretation

It presents the first agentic investigation environment with multiple adversarial interpretations grounded in real cases, using the court's opinion as judicially validated ground truth. Existing legal benchmarks test question answering, precedent retrieval, or bar-exam-style reasoning and assume a single correct answer or a well-posed query; InvestigationWorlds lets the same evidence support multiple coherent but incompatible conclusions and scores against the court-adopted hypothesis. The 100 cases span 28 states, the District of Columbia, and Puerto Rico across 41 federal judicial districts in all 12 regional circuits, balanced 20-per-category across 5 federal case types, with 42,283 pages of archived PACER PDFs; an attorney investigated 10 randomly selected cases with the opinion and adopted label withheld, reading 459 pages over 12 hours, and their factual conclusions agreed with the court-adopted outcome in all 10 cases.

It provides an attorney-validated generation pipeline that uses role-tagged documents (signal, noise, chaff, pro-alternative signal, alternative-specific noise) to build large corpora with controllable ambiguity around real records. Unlike end-to-end synthesized corpora or distractors set by retrieval similarity, the pipeline anchors on real case records, generates documents by evidentiary role against a hypothesis DAG, and releases all prompts and the role taxonomy. Three practicing attorneys with 30+ years of combined experience completed approximately 100 hours of expert review; in the 55-case stage-by-stage review, pass rates exceeded 94% for party hypotheses, alternative hypotheses, and chaff documents, with an 81%+ agreement score on a 10-case agreement subset; all 100 cases received a hypothesis-fidelity audit.

Evaluating six frontier agents on 100 cases, it finds the best model solves only one-third of cases and that retrieving relevant evidence does not prevent committing to an incorrect hypothesis. The result exposes a failure mode not captured by existing evaluations: agents may retrieve relevant evidence yet still believe an incorrect alternative hypothesis; it also shows a tradeoff between commitment frequency and conditional accuracy. Haiku 4.5 achieves the highest case-solved rate (33.3%), followed by Sonnet 4.6 (29.9%); all six models exceed the uniform-random baseline of 9.9%, but only these two exceed the stronger 25.0% baseline of believing one hypothesis at random and rejecting the rest, with gains of 8.3 and 4.9 percentage points that are not statistically significant; on the 87 cases with complete outputs from both models, Haiku commits on 80.7% of verdicts versus 73.3% for Sonnet, while Sonnet is more accurate among committed verdicts (74.1% versus 70.8%).

Through document-role ablations and a generator-swap experiment, it shows that pro-alternative evidence induces agents to change their judgments while evaluator rankings remain stable across generation pipelines. The ablation quantifies how corpus composition affects judgments, and the generator-swap experiment addresses whether the environment merely measures evaluator-generator alignment. In a cumulative ablation on 50 cases, adding 20 pro-alternative documents reduces case-solved rates for five of six models by 4 to 22 percentage points; across 506 paired verdicts, belief in an alternative hypothesis rises from 3/506 (0.6%) to 115/506 (22.7%), with 55 switches from do-not-believe and 59 from do-not-know; after rebuilding 30 held-out cases end-to-end with OpenAI models, evaluator rankings remain consistent up to ties, with no pair of evaluators reversing its strict ordering.

Perspective

The environments are drawn from U.S. federal civil cases that reached a fully briefed summary judgment motion, a posture that provides exactly the structured ground truth that makes rigorous evaluation possible; the authors see the methodology as generalizable to state-court records, administrative adjudications, and arbitration awards with similar evidentiary structures. For readers, it applies to research and engineering settings that need to evaluate whether an agent can construct and hold a correct factual narrative when there is no well-posed query and multiple pieces of evidence support several coherent accounts; the released PACER corpus, generation pipeline and prompts, and a portion of anonymized synthetic documents let other teams expand the case set or migrate to new domains.

The synthetic documents are judged plausible and role-consistent by experienced litigators but are not court-tested, and forensic-style realism (metadata patterns, threading conventions, native-format artifacts) can still be layered on incrementally; the current instantiation is limited to U.S. federal summary judgment cases, and transfer to other jurisdictions and adjudication types remains to be verified. In addition, the alternative hypotheses are generated by language models and audited by attorneys, so whether their difficulty distribution matches distractors in real investigations is an open question; the generator-swap experiment covers only 30 held-out cases, and while evaluator rankings remain consistent up to ties, individual models such as DeepSeek v4 show a large change in case-solved rate across generation conditions, whose stability warrants further observation.

Sources