PathReportEval: A Systematic Benchmark for Pathology Report Generation
Quick Answer
This paper shows that The PathReportEval framework standardizes pathology report generation evaluation, introducing the Clinical Report Quality Score (CRQS) to assess factual correctness.
Quick Take
It benchmarks four methods across three datasets (TCGA, HistAI, REG 2025) using encoders like CONCHv1.5 and UNI2-h, revealing significant discrepancies in clinical accuracy compared to traditional metrics.
Key Points
- Introduces a standardized benchmark for pathology report generation evaluation.
- Evaluates four methods using datasets TCGA, HistAI, and REG 2025.
- CRQS measures clinical fact coverage and hallucination rates effectively.
- Traditional metrics like BLEU and ROUGE often misrepresent report quality.
- Framework allows integration of new methods and datasets for future research.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Pathology report generation from whole-slide images (WSIs) is a rapidly growing multimodal learning problem, yet progress is difficult to measure because existing studies use heterogeneous datasets, model settings, visual encoders, and evaluation protocols. Moreover, commonly used natural language generation metrics, including BLEU, ROUGE, and METEOR, primarily reward lexical similarity and often fail to detect clinically consequential errors such as omitted diagnoses, hallucinated findings, or discordant tumor attributes.
We present a standardized benchmark and evaluation framework for pathology report generation. The benchmark evaluates four representative methods across three datasets (TCGA, HistAI, and REG 2025) using three pathology foundation encoders (CONCHv1.5, UNI2-h, and H-Optimus-1). Our framework standardizes preprocessing, feature extraction, training, decoding, and evaluation, enabling fair comparison across models while providing a modular platform for integrating new methods, datasets, and encoders.
A central contribution is the Clinical Report Quality Score (CRQS), a clinically grounded metric for evaluating factual correctness. CRQS maps reference and generated reports into structured clinical attributes and measures four complementary dimensions: clinical fact coverage, key information recall, hallucination rate, and clinical discordance, producing both an overall score and interpretable sub-scores.
Experiments demonstrate that conventional language-generation metrics are weakly aligned with clinical correctness and frequently overestimate report quality. In contrast, CRQS reveals clinically meaningful differences between models and encoders that lexical metrics fail to capture. Together, the benchmark, public plug-and-play framework, and CRQS establish a reproducible foundation for rigorous evaluation of pathology report generation.
| Subjects: | Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.18448 [cs.CL] |
| (or arXiv:2607.18448v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.18448 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Suryakant Singh [view email]
[v1]
Mon, 20 Jul 2026 18:59:47 UTC (1,022 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.