Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records
Quick Answer
This study employs a two-stage LLM pipeline (Gemini 2.5 Pro and Gemini 2.5 Flash) to identify 3,460 documentation inconsistencies in 3,000 MIMIC-IV-Note discharge summaries, affecting 69.7% of admissions.
Quick Take
The findings highlight the context-dependent nature of inconsistency detection and propose a graded ontology for better classification.
Key Points
- 3,460 inconsistencies identified, impacting 69.7% of discharge summaries.
- Inconsistencies span demographics, allergies, procedures, and more.
- Expert review revealed limitations in temporal reasoning and outpatient knowledge.
- Proposed a graded ontology for categorizing inconsistencies.
- Study lays groundwork for future large-scale EHR analysis.
DeepSignal Analysis
What happened
A study utilized a two-stage LLM pipeline to identify documentation inconsistencies in discharge summaries from the MIMIC-IV-Note dataset. The analysis revealed 3,460 inconsistencies across 3,000 summaries, impacting 69.7% of admissions. The findings indicate that inconsistency detection is context-dependent and propose a new ontology for classification.
Key evidence
- The study identified 3,460 candidate inconsistencies in 3,000 MIMIC-IV-Note discharge summaries, affecting 69.7% of admissions.
- The two-stage pipeline involved candidate identification using Gemini 2.5 Pro and context-grounded verification with Gemini 2.5 Flash.
- Expert review highlighted recurring failure modes, particularly in cases requiring temporal reasoning or knowledge of outpatient-prescribing conventions.
Why it matters
The findings underscore the challenges in reliably detecting inconsistencies in electronic health records (EHRs) using LLMs. The context-dependent nature of detection suggests that current models may struggle with certain types of inconsistencies, which could have implications for clinical decision-making and patient safety. Establishing a graded ontology for classification may enhance future analyses and improve the reliability of automated systems.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world discharge summaries, and to identify recurring failure modes that limit reliability at scale.
Materials and Methods: We applied a two-stage LLM pipeline---open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gemini 2.5 Flash)---to 3,000 randomly sampled MIMIC-IV-Note discharge summaries. A subset of the pipeline output was then reviewed manually by clinical experts.
Results: Our pipeline surfaced 3,460 candidate inconsistencies, affecting 69.7% of admissions. Representative examples spanned demographics, allergies, procedures, diagnoses, laboratory, medications, and care-planning domains, with direct implications for clinical reasoning or patient safety. Expert review also revealed recurring failure modes that arise when verification requires temporal reasoning, evolving-diagnosis context, or knowledge of outpatient-prescribing conventions the model does not natively possess.
Discussion: Detection is highly context-dependent: many flagged pairs require anchoring each statement to its source section and clinical domain, then assessing whether the conflict reflects a true contradiction or missing context. We propose a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis.
Conclusion: This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis.
| Subjects: | Computation and Language (cs.CL); Applications (stat.AP) |
| Cite as: | arXiv:2607.22954 [cs.CL] |
| (or arXiv:2607.22954v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22954 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Anru R. Zhang [view email]
[v1]
Fri, 24 Jul 2026 23:43:35 UTC (99 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.