Public articles linked to the same research event.
arXiv The authors introduce InvestigationWorlds: they retrieve 100 real U.S. federal civil cases with decided summary judgment motions from PACER and use an attorney-validated generation pipeline to synthesize role-tagged documents around the original record, so that one corpus admits multiple coherent but incompatible factual readings of which only one matches the court-adopted hypothesis; evaluating six frontier agents on 100 cases, they find agents often commit to incorrect hypotheses despite retrieving relevant evidence, with the highest case-solved rate only 33.3%.
The authors introduce InvestigationWorlds: they retrieve 100 real U.S. federal civil cases with decided summary judgment motions from PACER and use an attorney-validated generation pipeline to synthesize role-tagged documents around the original record, so that one corpus admits multiple coherent but incompatible factual readings of which only one matches the court-adopted hypothesis; evaluating six frontier agents on 100 cases, they find agents often commit to incorrect hypotheses despite retrieving relevant evidence, with the highest case-solved rate only 33.3%.
The authors introduce InvestigationWorlds: they retrieve 100 real U.S. federal civil cases with decided summary judgment motions from PACER and use an attorney-validated generation pipeline to synthesize role-tagged documents around the original record, so that one corpus admits multiple coherent but incompatible factual readings of which only one matches the court-adopted hypothesis; evaluating six frontier agents on 100 cases, they find agents often commit to incorrect hypotheses despite retrieving relevant evidence, with the highest case-solved rate only 33.3%.
The authors introduce InvestigationWorlds: they retrieve 100 real U.S. federal civil cases with decided summary judgment motions from PACER and use an attorney-validated generation pipeline to synthesize role-tagged documents around the original record, so that one corpus admits multiple coherent but incompatible factual readings of which only one matches the court-adopted hypothesis; evaluating six frontier agents on 100 cases, they find agents often commit to incorrect hypotheses despite retrieving relevant evidence, with the highest case-solved rate only 33.3%.