Schema-Constrained Document-Level Event Argument Extraction with Lightweight LLM Fine-Tuning
Quick Answer
This study demonstrates that mid-sized open LLMs, particularly Phi-4 (14B), can effectively perform schema-constrained Event Argument Extraction (EAE) at the document level, achieving a 42.39% F1 score on the MAVEN-ARG benchmark.
Quick Take
The approach utilizes role-set injection, LoRA fine-tuning, and deterministic decoding to enhance performance and reliability in extracting structured event records.
Key Points
- Phi-4 (14B) outperforms previous GPT models in event-coreference evaluations.
- The method incorporates role-set injection for schema compliance in prompts.
- LoRA fine-tuning is employed for efficient parameter optimization.
- Deterministic decoding includes post-processing to validate JSON outputs.
- Code for reproducing experiments is publicly available.
DeepSignal Analysis
What happened
The study investigates the capabilities of mid-sized open LLMs, particularly Phi-4 (14B), in performing schema-constrained Event Argument Extraction (EAE) at the document level. The model achieved a 42.39% F1 score on the MAVEN-ARG benchmark, indicating its effectiveness in this task. The methodology includes role-set injection, LoRA fine-tuning, and deterministic decoding to improve extraction performance.
Key evidence
- The research focuses on schema-constrained Event Argument Extraction (EAE) using mid-sized open LLMs, specifically Phi-4 (14B).
- The model achieved a 42.39% F1 score on the MAVEN-ARG benchmark, outperforming previous GPT baselines.
- The approach employs role-set injection, LoRA fine-tuning, and deterministic decoding to enhance the reliability of event record extraction.
Why it matters
This study highlights the potential of mid-sized open LLMs in handling complex document-level tasks like EAE, which is crucial for applications in natural language processing. The ability to achieve competitive performance with a 14B parameter model suggests that smaller models can be viable alternatives to larger counterparts, potentially reducing resource requirements for similar tasks. This could democratize access to advanced NLP capabilities.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Event Argument Extraction (EAE) converts documents into structured event records by identifying argument spans and assigning them schema-defined roles. Document-level EAE is challenging due to long-range dependencies between triggers and arguments, cross-sentence context, and strict role constraints, which often lead to boundary errors, uncertainty in roles, and inconsistencies with restricted schemas.
In this paper, we study whether mid-sized open LLMs can perform schema-constrained EAE reliably at the document level on MAVEN-ARG. Our approach combines (i) role-set injection in prompts for schema compliance, (ii) parameter-efficient supervised fine-tuning (LoRA) using the same JSON-only interface used at inference, and (iii) deterministic decoding with post-processing that validates JSON, filters invalid roles, de-duplicates arguments, and aligns spans to the document window. Under the official MAVEN-ARG evaluator, fine-tuned mid-sized open models outperform previously reported GPT baselines across mention, entity-coreference, and event-coreference evaluations; our best model (Phi-4, 14B) reaches 42.39\% F1 at the event-coreference level. Code to reproduce experiments is publicly available at this https URL.
| Comments: | Accepted at ECML PKDD 2026 |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.16808 [cs.CL] |
| (or arXiv:2607.16808v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16808 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Antonio Guerriero [view email]
[v1]
Sat, 18 Jul 2026 13:04:06 UTC (518 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.