What Carries the Signal in Pathology Foundation-Model Atlases? A Patient-Level Controlled Benchmark in Breast Cancer
Quick Answer
This study evaluates pathology foundation models in breast cancer, revealing that ridge regression on mean-pooled embeddings predicts held-out program scores with Spearman rho up to 0.556.
Quick Take
The research indicates that while the signal is real, it is not uniformly morphological, with embeddings outperforming tissue composition in several gene programs.
Key Points
- Ridge regression achieved Spearman rho scores between 0.25 and 0.56 across 285 TCGA-BRCA patients.
- Embeddings outperformed tissue composition in predicting immune, proliferation, and ER/luminal scores.
- Geometric machinery provided no measurable contribution to model performance.
- Fifty-four interpretable cell-count features closely matched program predictions.
- Driver-count metrics were largely uninformative, with 91.8% of random panels recovering at least 5 drivers.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Pathology foundation models are reported to encode molecular programmes in tissue morphology, but the evidence is usually a cohort-wide ranked gene list rather than a prediction for a held-out patient. We rebuild such an analysis with the patient as the unit of evidence and ask which pipeline component carries signal.
Across 11 frozen backbones, four pre-specified gene programmes and 285 TCGA-BRCA patients with paired slides and RNA-seq (44 cells; GroupKFold by patient, all preprocessing fitted inside the fold), ridge regression on mean-pooled embeddings predicts held-out programme scores at Spearman rho = 0.25-0.56, UNI2 strongest on all four (immune 0.556). A matched permutation null gives raw p ~ 1e-4 at 10,000 permutations for every cell; Holm-adjusted p = 0.0044.
The signal is real but not uniformly morphological. Against competing models on the same patients and folds, embeddings beat tissue composition for ER/luminal, proliferation and immune (+0.280, +0.284, +0.479; p <= 0.003) but not basal, where compartment fractions alone reach 0.469 against the embedding's 0.493 (p = 0.77). Fifty-four interpretable cell-count features come within 0.043-0.085 on every programme.
The geometric machinery contributes nothing measurable, and we identify why: the geodesic graph selects neighbours by Euclidean nearest-neighbour search and only reweights edges already chosen, so the topology is Euclidean by construction (Riemannian minus Euclidean = +0.0010, 95% CI [-0.0007, +0.0029]). Applied consistently the geometry is worse (-0.0117). Ridge regression beats the graph-and-metric decoder by +0.097 (CI [+0.069, +0.127]). The driver-count metric common in this literature is near-uninformative here: 91.8% of random six-gene panels recover >=5/6 drivers.
| Comments: | 40 pages, 7 figures, 10 tables. Supplementary Information (16 pages) included as an ancillary file. Segmentation outputs obtained under the Aignostics Research Access Programme; OpenTME data at this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Quantitative Methods (q-bio.QM) |
| Cite as: | arXiv:2608.00105 [cs.CV] |
| (or arXiv:2608.00105v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.00105 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chimdi Walter Ndubuisi [view email]
[v1]
Fri, 31 Jul 2026 04:58:29 UTC (7,587 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.