Skip to main content
Back to timeline
arXivSource publication:

LexReward scores Chinese legal responses along Style, Element and Chain, and its reward models lift GRPO's Element average by 6.04 points over SFT

Related research and updates

Synopsis

The work introduces LexReward, a taxonomy built with legal experts that decomposes legal response quality into Style (lexical and syntactic), Element (subjects, facts, statutes, decisions) and Chain (order, completeness, correctness, non-redundancy), designs rubrics scored by rules or an LLM judge, uses the resulting rewards to build preference data and train LexRM, a family of Chinese legal reward models; experiments show the rubrics distinguish responses of different quality, DPO improves generation across all three dimensions, LexRM-guided selection beats random selection in test-time scaling, and LexRM used for GRPO yields gains in every dimension and outperforms rule-based outcome rewards.

AI-generated editorial illustration: LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models

Interpretation

It proposes a three-dimension taxonomy of legal response quality, decomposing the requirements of a high-quality legal response into Style, Element and Chain with fine-grained criteria. Existing legal reward approaches often rely on holistic judgments or assess only selected aspects such as judgment-outcome accuracy, leaving the definition of legal response quality unstructured; this work combines top-down expert domain knowledge with bottom-up error analysis to make those requirements explicit and assessable. Taxonomy construction is described in Section 3, involving legal experts, comparison of multiple models' responses with reference answers across legal AI benchmarks, and a dimension ablation (Table 3) in which aggregation gives the best overall score in all three dimensions.

It designs interpretable rubrics for each fine-grained criterion, using rule-based scoring or an LLM judge depending on whether the criterion requires input context, and shows these rewards reliably distinguish responses of different quality. Rather than holistic preferences or a single outcome reward, the rewards attach to specific criteria: Style is scored by rule-based functions calibrated against authentic judicial documents, while Element and Chain use an LLM judge or rule-based scoring depending on the task. Table 1 shows the rubric outperforms the five-model average and random selection on all three datasets, e.g. 79.00 versus 50.00 accuracy on CLASE for Style; Table 3 shows single sub-dimensions score below their aggregation overall.

It trains LexRM, described by the authors as the first family of reward models for the Chinese legal context, from rubric-derived preference data, and LexRM-guided selection outperforms random selection in test-time scaling. Because rubrics can assign identical scores and create ties, the work trains continuous scoring functions from preference pairs to separate candidates the rubrics cannot, trains dimension-specific reward models, and merges them into one multi-dimensional model via task arithmetic. In Table 1, LexRM-Style reaches 81.75 on CLASE and LexRM-Chain reaches 50.46 overall on LexChain, the best results on their respective datasets, while LexRM-Element ranks second on Legal; LexRM-Merge matches the element expert on Legal and obtains the highest charge-prediction score but falls below the respective expert models on CLASE and LexChain.

Rubric-derived preference data improve generation across all three dimensions through DPO, and LexRM used for GRPO yields gains in every dimension and outperforms rule-based outcome rewards. Legal RL is currently driven almost entirely by outcome rewards defined on a verifiable final answer; this work shows dimensional rewards supervise the reasoning that leads to an answer rather than only estimating the answer itself. In Table 2, DPO improves over the vanilla model on all three dimensions; GRPO (RM) achieves the highest score on every dimension, improving all five Element metrics and raising their average by 6.04 points over SFT; GRPO (rule) gains are confined to the final answer while Element and Chain averages remain largely unchanged.

Perspective

The framework targets the Chinese legal context: the taxonomy and rubrics are built from Chinese legal materials and evaluation is limited to Chinese-language datasets, so its conclusions apply to generating and optimizing Chinese legal responses. It lets legal domain knowledge directly guide reward construction and model optimization: rubrics can distinguish response quality, preference data can be used for DPO, and LexRM can select among candidates in test-time scaling or serve as a GRPO reward to optimize Style, Element and Chain separately. For legal AI researchers and engineers who want dense, interpretable reward signals without reference answers, this offers a reusable construction pipeline.

The taxonomy and evaluation are limited to Chinese legal materials, and applicability to other languages and legal systems remains untested. The work primarily studies dimension-specific rewards; combining multiple dimensions into a unified rubric reward or reward model remains an open challenge, including aggregation weights, preference-data composition and training strategy. Rubrics may also assign identical scores to responses of differing quality, and queries where all candidates receive identical scores are discarded during preference construction, so how this affects the coverage of the resulting reward models is worth watching in follow-up work. This is a full-text reading, but figures are presented as text and some implementation details and prompt templates live in the appendices, so reproduction should combine the appendices with the released data and code.

Sources