CONSISTRE adds consistency constraints and distilled RL so black-box and 7–8B open models extract relations more consistently on DocRED
Synopsis
The work proposes CONSISTRE, a unified consistency-aware framework for document-level relation extraction (DocRE) with two complementary tracks: an inference-time track for black-box LLMs that combines constraint-aware prompting, constraint-based verification, and iterative self-reflection without task-specific fine-tuning, and a training-time track that injects consistency knowledge into smaller open-source models by distilling reasoning traces from a powerful teacher via supervised fine-tuning followed by GRPO alignment with a composite reward jointly optimizing extraction performance and relational consistency; on DocRED both tracks outperform their baselines, with the inference-time track achieving competitive F1 using off-the-shelf black-box LLMs and the training-time track substantia
Fig. 2: Per-type Cons across the three stages of Track B on the 200-document DocRED evaluation subset. SFT primarily contributes to inverse-relation Cons; GRPO primarily contributes to transitivity Cons (Qwen3-8B: +0.181; Qwen2.5-7B: +0.235). See Section IV-G for discussion.
arXivInterpretation
It proposes CONSISTRE, a unified consistency-aware DocRE framework that explicitly incorporates relational constraints such as transitivity, symmetry, and functional uniqueness into the extraction pipeline, covering both API-accessible and locally deployable scenarios. Prior LLM-based information extraction typically generates predictions independently for each candidate triple, which can violate fundamental relational constraints and produce contradictory outputs; this work places both deployment paradigms under a unified consistency formulation. At the abstract level it presents the framework design and the composition of the two tracks; the specific constraint forms and the details of the unified formulation are not expanded in the abstract.
The inference-time track targets black-box LLMs, combining constraint-aware prompting, constraint-based verification, and iterative self-reflection to refine predictions without task-specific fine-tuning. It moves consistency handling to inference rather than training, so off-the-shelf black-box models whose weights are inaccessible or inconvenient to fine-tune can also benefit. On DocRED this track outperforms its baselines and achieves competitive F1 using off-the-shelf black-box LLMs; the specific baseline settings and numbers are not given in the abstract.
The training-time track injects consistency knowledge into smaller open-source models via knowledge distillation and reinforcement learning: reasoning traces from a powerful teacher are distilled into a student via supervised fine-tuning, followed by GRPO alignment with a composite reward that jointly optimizes extraction performance and relational consistency. It turns consistency from inference-time prompting and verification into a trainable objective signal, giving locally deployable small models consistency capability. On DocRED this track substantially narrows the gap between 7–8B open-source models and state-of-the-art proprietary LLMs at a fraction of their inference cost; the specific reward weights and training details are not given in the abstract.
Ablation studies confirm that explicit consistency modeling mitigates relational contradictions and enhances the reliability of LLM-based DocRE across both deployment paradigms. It separates the contribution of consistency modeling from overall performance, indicating the gains do not come only from stronger extraction ability. The abstract reports ablation studies supporting this conclusion, but does not give the specific ablation configurations or numbers.
Perspective
The work targets document-level relation extraction, especially settings that require maintaining triple consistency across sentences and entities; its inference-time track suits users who can only access black-box LLMs via API, and its training-time track suits teams that want to deploy 7–8B open-source models locally and control inference cost. The experiments described in the abstract center on the DocRED dataset, so the conclusions mainly apply to the document-level relation extraction setting that benchmark represents; the consistency constraints are exemplified by relational properties such as transitivity, symmetry, and functional uniqueness, so they apply to relation types that have such structural constraints.
The abstract does not give specific F1 values, baseline configurations, the sizes of the teacher and student models, the composition and weights of the composite reward, or the specific ablation settings, so the magnitude of the consistency gain relative to extraction performance improvement cannot be judged. How constraint verification and iterative self-reflection perform on longer documents or more complex relation types, and the stability of GRPO alignment, remain questions a reader would want to confirm. Because this is based only on the abstract, these details are unexpanded parts and the original text should be consulted.
