Public articles linked to the same research event.
arXiv The work introduces JEVal, a bilingual decision benchmark of 11,257 instances from 36 datasets across 10 domains, evaluates 25 configurations and finds that general decision models approach strong thinking LLMs when evidence is available but weaken on specialist knowledge, calibrated uncertainty, exact computation, and long-context evidence use, with local decision advantages offset by system-level errors in long-horizon agents and social simulation; the authors then build InnerJev-4B and InnerJev-27B from open-weight LLMs via first-token decision readout and Reasoning-to-Readout Self-Distillation, with InnerJev-27B reaching 79.24% on JEVal versus Jev's 78.38% and answering a typical query in about 0.1 s.
The work introduces JEVal, a bilingual decision benchmark of 11,257 instances from 36 datasets across 10 domains, evaluates 25 configurations and finds that general decision models approach strong thinking LLMs when evidence is available but weaken on specialist knowledge, calibrated uncertainty, exact computation, and long-context evidence use, with local decision advantages offset by system-level errors in long-horizon agents and social simulation; the authors then build InnerJev-4B and InnerJev-27B from open-weight LLMs via first-token decision readout and Reasoning-to-Readout Self-Distillation, with InnerJev-27B reaching 79.24% on JEVal versus Jev's 78.38% and answering a typical query in about 0.1 s.
The work introduces JEVal, a bilingual decision benchmark of 11,257 instances from 36 datasets across 10 domains, evaluates 25 configurations and finds that general decision models approach strong thinking LLMs when evidence is available but weaken on specialist knowledge, calibrated uncertainty, exact computation, and long-context evidence use, with local decision advantages offset by system-level errors in long-horizon agents and social simulation; the authors then build InnerJev-4B and InnerJev-27B from open-weight LLMs via first-token decision readout and Reasoning-to-Readout Self-Distillation, with InnerJev-27B reaching 79.24% on JEVal versus Jev's 78.38% and answering a typical query in about 0.1 s.
The work introduces JEVal, a bilingual decision benchmark of 11,257 instances from 36 datasets across 10 domains, evaluates 25 configurations and finds that general decision models approach strong thinking LLMs when evidence is available but weaken on specialist knowledge, calibrated uncertainty, exact computation, and long-context evidence use, with local decision advantages offset by system-level errors in long-horizon agents and social simulation; the authors then build InnerJev-4B and InnerJev-27B from open-weight LLMs via first-token decision readout and Reasoning-to-Readout Self-Distillation, with InnerJev-27B reaching 79.24% on JEVal versus Jev's 78.38% and answering a typical query in about 0.1 s.