LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction
Quick Answer
LA-RL introduces a label-aware self-reflection framework for information extraction, enhancing performance in tasks like named entity recognition and relation extraction.
Quick Take
The model achieves an average F1 score of 6.83 on SciER relation extraction and shows significant improvements over standard fine-tuning methods, particularly in out-of-distribution scenarios.
Key Points
- LA-RL uses task-grounded diagnostic labels for self-correction in information extraction.
- Achieved 6.83 average F1 score on SciER relation extraction benchmarks.
- Significant gains observed in out-of-distribution relation extraction tasks.
- Training involves cold-start supervised fine-tuning and two stages.
- Reflection structure is task-sensitive, benefiting relation extraction more than named entity recognition.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models show strong promise for information extraction (IE), but existing reflection-based correction methods are often misaligned with structured extraction outputs. Free-form self-reflection can flag an error, yet it rarely identifies whether the failure is a missing span, wrong label, boundary mismatch, invalid relation type, or reversed argument order. We introduce LA-RL (Label-Aware Reflective Reinforcement Learning), an outcome-supervised framework that guides IE self-correction with task-grounded diagnostic labels. A single backbone first predicts an extraction, diagnoses task-specific error labels, and then revises its output conditioned on the diagnosis. Training starts from diagnostic data labeled by an annotation model for cold-start supervised fine-tuning and proceeds through two GRPO stages that reward final extraction quality, format validity, and first-pass correctness, without a process reward model. Experiments on named entity recognition, relation extraction, and event extraction show consistent same-backbone gains over SFT, including 6.83 average F1 on SciER relation extraction, about 20 F1 on out-of-distribution relation extraction, and 14.80 trigger F1 plus 17.50 argument F1 on DuEE1.0. Ablations show that reflection structure is task-sensitive: stronger constraints benefit relation extraction, whereas named entity recognition needs less restrictive correction under domain shift.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.23420 [cs.CL] |
| (or arXiv:2607.23420v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.23420 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xiao You [view email]
[v1]
Sun, 26 Jul 2026 02:35:47 UTC (392 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.