Fudan team trains Chinese-Jev on 10 million Chinese decisions, reaching 69.20% general accuracy and roughly 20x faster than Jev
Synopsis
The work introduces Chinese-Jev: a unified data pipeline converts heterogeneous Chinese annotations into probability targets over candidate options, a lightweight encoder-only model is first pre-trained on 10 million general Chinese decisions and then fine-tuned separately for medicine, law, and finance, and CJ-Bench is released with 307,900 held-out decisions; after pre-training it reaches 69.20% accuracy on general tasks, a 1.24% relative gain over the closed-source Jev, cuts expected calibration error from 11.45% to 3.78% at 14 ms latency, improves medical accuracy over Jev by 4.0%, reaches 92% of Jev's average accuracy across the three specialized domains at 15 ms average latency, and an INT8-quantized model runs about 1 second per decision in a mobile browser.
Interpretation
It presents Chinese-Jev, a System One model for Chinese decision-making that supports choice, noul (proposition judgment), and score (ordered rating) decisions, with a general model and three domain specialists sharing one architecture and decision interface that outputs candidate probabilities in a single forward pass. Prior Jev-style models were mainly English-oriented, and existing Chinese releases focused on football-domain questions or bilingual encoders; this work extends the unified decision interface to general Chinese tasks across eight categories and adapts it independently to three specialized domains. The paper details the architecture: an mmBERT-base backbone (22 layers, hidden size 768) plus two decision Transformer layers and a shared candidate scorer, about 322 million parameters, fully fine-tuned on eight NVIDIA H200 GPUs.
It designs a reusable data construction pipeline that converts single-answer questions and category labels, true/false and proposition annotations, discrete and continuous ratings, and multiple-answer questions into probability distributions over candidate options, with decision-level deduplication and type budgets, yielding 10 million general Chinese decisions (roughly one third each for choice, noul, and score) plus 3 million medical and 24,000 each for legal and finance. Existing Chinese resources such as CLUE, C-Eval, CMMLU, and COIG-CQIA provide classification labels, instruction data, or evaluation suites, whereas this pipeline unifies categorical labels, proposition judgments, and ordered ratings into one candidate-scoring form and keeps soft targets that preserve mean ratings and inter-rater variation for continuous ratings. The paper lists contribution shares from 20 public sources (for example T2Ranking at 23.28% and ASAP at 13.19%) plus 10,700 rule-based synthetic decisions, and assigns decisions derived from the same source material to the same data partition to reduce leakage.
It releases CJ-Bench with 307,900 held-out decisions (100,000 general, 200,000 medical, 4,300 legal, 3,600 finance) to evaluate accuracy, calibration, and latency under a common protocol. Earlier Chinese evaluations largely target knowledge and reasoning of generative models, while this benchmark places decision accuracy, expected calibration error, and end-to-end latency under one protocol spanning general and three specialized domains. The general component is sampled without replacement from the test pool with roughly equal choice, noul, and score budgets, each source contributing at most 15%, stratified by candidate count, label, and input length; ECE uses 15 equal-width bins.
Experiments show 69.20% general accuracy after pre-training, 0.85 percentage points above the closed-source Jev (1.24% relative), with ECE dropping from 11.45% to 3.78% at 14 ms latency; after domain fine-tuning, medical accuracy is 87.65%, 3.36 points above Jev (4.0% relative), legal is 60.10% and finance 64.72%, averaging 92% of Jev across the three domains at 15 ms average latency; the INT8-quantized model runs about 1 second per decision on an iPhone 15 Pro. Relative to the multilingual initialization Laya Multilingual, Chinese pre-training raises accuracy from 40.22% to 69.20% (+28.98 points) and lowers ECE from 23.04% to 3.78%; domain fine-tuning adds 47.88, 21.34, and 29.33 points over the general model on medical, legal, and finance respectively. Results come from CJ-Bench components with all models receiving the same retained context and candidates; latency is median end-to-end time on one H200 at batch size one after warm-up, and the comparison with the hosted Jev API includes network round-trip time, which the paper notes.
Perspective
The work targets Chinese applications that need to choose among candidates, judge whether a proposition holds, or assign an ordered rating, such as request routing, relevance judgment, and answer selection; the general model covers eight Chinese decision task categories, while the medical, legal, and finance specialists apply to their respective domains and share one interface so a domain model can be swapped without changing how decisions are specified. The INT8-quantized version targets local inference in a mobile browser, keeping input text on the device. The paper states it will release models, training data, the benchmark, and data construction code, so follow-up work can extend the same pipeline to new Chinese tasks or domains and continue comparing accuracy, calibration, and latency on CJ-Bench.
The legal and finance specialists still trail Jev in accuracy, which the paper lists as motivation for further work; calibration gains after domain adaptation are uneven, with legal ECE rising from 15.37% to 17.45% and the specialists' average ECE at 12.08% versus Jev's 5.57%. The latency comparison with the hosted Jev API includes network round-trip time, which the paper notes reflects end-to-end response time rather than pure model inference speed. The near-duplicate screen covers only held-out text available when a candidate is checked, so paraphrases across the full training corpus are not exhaustively merged. The mobile result comes from a browser prototype on one device and OS version, leaving open whether the roughly 1-second latency holds on other hardware, and how excluding soft-target score decisions from accuracy and calibration statistics affects the picture.
