Skip to main content
Back to timeline
arXivSource publication:

Researchers release the bilingual JEVal decision benchmark and train InnerJev-27B, which reaches 79.24% accuracy on 11,257 items versus Jev's 78.38% at about 0.1 s per query

Related research and updates

Synopsis

The work introduces JEVal, a bilingual decision benchmark of 11,257 instances from 36 datasets across 10 domains, evaluates 25 configurations and finds that general decision models approach strong thinking LLMs when evidence is available but weaken on specialist knowledge, calibrated uncertainty, exact computation, and long-context evidence use, with local decision advantages offset by system-level errors in long-horizon agents and social simulation; the authors then build InnerJev-4B and InnerJev-27B from open-weight LLMs via first-token decision readout and Reasoning-to-Readout Self-Distillation, with InnerJev-27B reaching 79.24% on JEVal versus Jev's 78.38% and answering a typical query in about 0.1 s.

Source-provided article image: General Decision Models: Benchmarking and Insights Beyond Jev
Figure 1 ·

Figure 1: A conceptual view of general and decisional intelligence. The left panel contrasts open-ended language capabilities with decision-making in constrained output spaces. The right panel illustrates the evolution from expert systems and task-specific models to encoder models, chat and reasoning LLMs, and general decision models represented by Jev.

arXiv

Interpretation

Introduces JEVal, a bilingual benchmark that unifies diverse tasks into a finite decision space with three formats (choice, noul, score), comprising 11,257 instances from 36 datasets, 58 subtasks, and 10 domains, with 9,457 English and 1,800 Chinese instances. Prior evaluations of decision models focused on individual tasks or single reliability aspects; JEVal places knowledge reasoning, long-context understanding, agentic action selection, medicine, law, finance, and personalization under one decision interface with four metrics: accuracy, ECE, raw probability-vector squared error, and latency. Benchmark construction involved two human experts inspecting task definitions and sample content, removing instances unsuitable for decision-task formulation, and retaining only datasets whose reference answers were manually verified; after exact-string deduplication from over 202K candidates, at most 200 instances per subtask were resampled, with input length controlled against a 20,480-token budget and candidate options capped at 255.

Systematically characterizes the capability boundaries of general decision models: they are most competitive when evidence is contained in the input, but weaken when decisions require specialist knowledge, faithful uncertainty estimation, or exact computation. Prior evaluations of Jev-like models largely stayed at the single-task level; this work separates choosing the most likely outcome from estimating its probability using JEVal plus four controlled diagnostic tasks of 100 items each. On JEVal, the strongest decision models fall within three points of leading thinking LLMs in overall accuracy; in medicine, finance, and law, the best decision model is three to six points below the best generative LLM. In probability tasks, Jev picks the mode on all 100 prior items but assigns the true mode 89.35% of the mass on average against 35.60% in the true distribution, and its two highest-probability posterior options receive 90.30% of the mass. On repeating-decimal division and area inversion, Jev answers 44% and 39% exactly, yet 97% and 73% of its answers fall within one integer of the gold value.

Moves evaluation from static benchmarks to dynamic systems: in the -bench tool-calling agent, using Jev for next-action selection cuts median episode time from 267.1 s to 156.1 s but lowers task success from 77.6% to 66.1% and raises P95 latency from 695.5 s to 1,009.5 s; in large-scale social simulation, decision models approach strong generative LLMs on individual prediction at far lower inference cost but show larger aggregate estimation errors and systematic bias. Prior work rarely examined local decision accuracy and system-level behavior together; this work uses a matched protocol (same user simulator, temperature 0, environments, and reward checker) to isolate the action-selection component, and drives about 330,000 voters in ElectionSim with InnerJev-27B. The agent experiments cover 115 Retail and 50 Airline tasks, yielding 330 interaction episodes with up to 30 turns each; failure trajectories show errors arising from inappropriate action selection or subsequent incorrect tool arguments, both triggering repeated lookups and retries. In social simulation, InnerJev-27B calls 47 of 51 states and 12 of 15 battleground states, one more of each than the original GPT-4o-mini pipeline, but its vote-share MAE is 6.52 versus 3.06 and its bias +3.73 versus +2.31. On individual survey prediction, InnerJev-27B reaches 39.87% overall accuracy, 0.54 points below the best configuration, at a mean latency of 8.36 s versus 211.65 s.

Proposes First-Token Decision Readout and Reasoning-to-Readout Self-Distillation, building InnerJev-4B and InnerJev-27B from open-weight LLMs, with InnerJev-27B reaching 79.24% on JEVal versus Jev's 78.38% and the lowest ECE and probability-vector squared error among all 25 configurations. Jev is available only through a hosted API and its training recipe is not public; this work shows that without labeled answers or a stronger teacher, distilling the answer distribution at the end of the model's own reasoning chain into its first answer token internalizes reasoning into a single forward pass, with embeddings and output layer frozen so inference cost matches the untrained readout. Training uses 39,279 questions, a single epoch, BF16 precision, and full-rank updates to attention and MLP projections (LoRA versions land within a point); on a single H100, median latency is about 58.7 ms for InnerJev-4B and 112.2 ms for InnerJev-27B. On newly constructed questions that played no role in training or model selection, InnerJev-27B improves on the untrained readout by about seven points.

Perspective

The results target researchers and engineering teams that need structured judgment and selection over finite output spaces: when decisions can be resolved directly from input evidence, general decision models and the InnerJev family can serve as low-latency, low-cost replacements for Jev; when specialist knowledge, exact numerical computation, or faithful probability estimation is required, thinking-mode generative LLMs or heterogeneous composition should be retained. The InnerJev readout covers choice, noul, and score formats, and training data only requires questions with a well-defined option set, so it transfers directly to label-selection tasks such as intent recognition, harmful-request judgment, and emotion classification. In social simulation, each voter costs two forward passes and no generated tokens, suiting macro-level simulations that need thousands of parallel agents.

Several points warrant attention: JEVal's finance domain contains only 59 instances, and accuracy on social-science GlobalOpinionQA and personalization tasks is generally low, so conclusions in those areas are sensitive to sample size and task difficulty; the agent experiments cover only Retail and Airline, and the authors note that the trajectory audit identified action errors and integration-level empty responses without isolating a dominant cause; in social simulation, InnerJev-27B's vote-share errors are about twice those of the original pipeline, and every run overestimates the Democratic share, a bias also reported for generative agents, indicating that moving from generation to decision preserves the lean; the training-set-size curves show gains stalling beyond 39,279 questions, which the authors attribute to self-written questions being easy and hard ones being hard to verify, so further scaling of self-distillation may depend on harder and more reliably posed question sources; finally, development-set checkpoint selection makes absolute scores optimistic, and the curves support the overall trend but not fitting a scaling law.

Sources