Public articles linked to the same research event.
arXiv The work introduces LexReward, a taxonomy built with legal experts that decomposes legal response quality into Style (lexical and syntactic), Element (subjects, facts, statutes, decisions) and Chain (order, completeness, correctness, non-redundancy), designs rubrics scored by rules or an LLM judge, uses the resulting rewards to build preference data and train LexRM, a family of Chinese legal reward models; experiments show the rubrics distinguish responses of different quality, DPO improves generation across all three dimensions, LexRM-guided selection beats random selection in test-time scaling, and LexRM used for GRPO yields gains in every dimension and outperforms rule-based outcome rewards.
The work introduces LexReward, a taxonomy built with legal experts that decomposes legal response quality into Style (lexical and syntactic), Element (subjects, facts, statutes, decisions) and Chain (order, completeness, correctness, non-redundancy), designs rubrics scored by rules or an LLM judge, uses the resulting rewards to build preference data and train LexRM, a family of Chinese legal reward models; experiments show the rubrics distinguish responses of different quality, DPO improves generation across all three dimensions, LexRM-guided selection beats random selection in test-time scaling, and LexRM used for GRPO yields gains in every dimension and outperforms rule-based outcome rewards.
The work introduces LexReward, a taxonomy built with legal experts that decomposes legal response quality into Style (lexical and syntactic), Element (subjects, facts, statutes, decisions) and Chain (order, completeness, correctness, non-redundancy), designs rubrics scored by rules or an LLM judge, uses the resulting rewards to build preference data and train LexRM, a family of Chinese legal reward models; experiments show the rubrics distinguish responses of different quality, DPO improves generation across all three dimensions, LexRM-guided selection beats random selection in test-time scaling, and LexRM used for GRPO yields gains in every dimension and outperforms rule-based outcome rewards.
The work introduces LexReward, a taxonomy built with legal experts that decomposes legal response quality into Style (lexical and syntactic), Element (subjects, facts, statutes, decisions) and Chain (order, completeness, correctness, non-redundancy), designs rubrics scored by rules or an LLM judge, uses the resulting rewards to build preference data and train LexRM, a family of Chinese legal reward models; experiments show the rubrics distinguish responses of different quality, DPO improves generation across all three dimensions, LexRM-guided selection beats random selection in test-time scaling, and LexRM used for GRPO yields gains in every dimension and outperforms rule-based outcome rewards.