Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

LexReward scores Chinese legal responses along Style, Element and Chain, and its reward models lift GRPO's Element average by 6.04 points over SFT

The work introduces LexReward, a taxonomy built with legal experts that decomposes legal response quality into Style (lexical and syntactic), Element (subjects, facts, statutes, decisions) and Chain (order, completeness, correctness, non-redundancy), designs rubrics scored by rules or an LLM judge, uses the resulting rewards to build preference data and train LexRM, a family of Chinese legal reward models; experiments show the rubrics distinguish responses of different quality, DPO improves generation across all three dimensions, LexRM-guided selection beats random selection in test-time scaling, and LexRM used for GRPO yields gains in every dimension and outperforms rule-based outcome rewards.