EAGER boosts generative event extraction with verifiable-reward reinforcement learning, beating prompting, supervised fine-tuning, and prior RL baselines across seven benchmarks
Synopsis
The work presents EAGER, a reinforcement learning framework for generative event extraction that combines fine-grained verifiable rewards with Schema-Contrastive Advantage Estimation to alleviate advantage collapse under sparse binary rewards; its reward design explicitly targets structural validity, extraction accuracy, groundedness, coverage, over-generation, and span precision, and across seven benchmark datasets it consistently outperforms prompting, supervised fine-tuning, and prior reinforcement learning baselines, achieving a substantial improvement over the strongest prior method.
Interpretation
EAGER casts generative event extraction as a reinforcement learning problem optimizable with verifiable rewards, with reward signals explicitly targeting structural validity, extraction accuracy, groundedness, coverage, over-generation, and span precision. Relative to prompting or supervised fine-tuning alone, the work brings task-aligned fine-grained verifiable rewards into end-to-end event extraction, so trigger identification, event type classification, and schema-grounded argument span extraction are directly constrained within one reward framework. The text supports this through the reward design description plus experiments on seven benchmarks, i.e., method description with benchmark comparison.
EAGER introduces Schema-Contrastive Advantage Estimation to alleviate advantage collapse under sparse binary rewards. Prior reinforcement learning baselines can suffer degraded advantage signals under sparse binary rewards; this work adds a contrastive advantage estimation mechanism as a targeted component. The text presents this mechanism as a core part of the framework and supports it with comparisons against prior reinforcement learning baselines.
Across seven benchmark datasets, EAGER consistently outperforms prompting, supervised fine-tuning, and prior reinforcement learning methods, with a substantial improvement over the strongest prior method. The result tests the combination of verifiable rewards and contrastive advantage estimation in a multi-dataset comparison rather than on a single dataset. Evidence comes from comparisons on seven benchmarks; the text summarizes this as 'consistently outperforms' and 'substantial improvement' without giving specific numbers in the abstract.
Perspective
The work targets end-to-end generative event extraction, suited to extraction settings that must jointly output triggers, event types, and schema-grounded argument spans, with seven benchmark datasets as the main validation setting. Its verifiable reward design targets structural validity, extraction accuracy, groundedness, coverage, over-generation, and span precision, so the most direct beneficiaries are extraction tasks with explicit schemas and automatically checkable outputs; the text gives no applicability statement for settings lacking verifiable signals or with unfixed schemas.
The text is abstract-level and provides no per-benchmark metrics, improvement magnitudes, statistical significance, weights for the reward components, or the concrete computation and ablation of Schema-Contrastive Advantage Estimation. Readers should therefore watch how large the 'substantial improvement' over the strongest prior method actually is, how much each of the six reward aspects contributes, and whether the alleviation of advantage collapse holds consistently across event types and schema complexities.
