Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Researchers release the bilingual JEVal decision benchmark and train InnerJev-27B, which reaches 79.24% accuracy on 11,257 items versus Jev's 78.38% at about 0.1 s per query

The work introduces JEVal, a bilingual decision benchmark of 11,257 instances from 36 datasets across 10 domains, evaluates 25 configurations and finds that general decision models approach strong thinking LLMs when evidence is available but weaken on specialist knowledge, calibrated uncertainty, exact computation, and long-context evidence use, with local decision advantages offset by system-level errors in long-horizon agents and social simulation; the authors then build InnerJev-4B and InnerJev-27B from open-weight LLMs via first-token decision readout and Reasoning-to-Readout Self-Distillation, with InnerJev-27B reaching 79.24% on JEVal versus Jev's 78.38% and answering a typical query in about 0.1 s.