Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining
Quick Answer
FlexiCT introduces a novel family of CT foundation models trained on 266,227 volumes, outperforming previous task-specific models in segmentation, classification, and more.
Quick Take
This agglomerative pretraining approach enhances CT representation learning, aligning imaging features with disease phenotypes across multiple benchmarks.
Key Points
- FlexiCT trained on 266,227 CT volumes from 56 datasets.
- Achieves superior performance across five task families including segmentation and classification.
- Utilizes three-stage agglomerative pretraining for enhanced CT representation.
- Embeddings organize CT scans by tumor stages, aiding disease phenotype characterization.
- Code available for public access to support further research.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Computed tomography (CT) is a central to three-dimensional medical imaging, yet CT-based artificial intelligence remains fragmented across task-specific models for segmentation, classification, registration, and report analysis. Here we present FlexiCT, a family of CT foundation models trained by agglomerative continual pretraining on 266,227 CT volumes from 56 publicly available datasets, forming a large-scale public resource for CT representation learning. FlexiCT uses agglomerative pretraining across three stages: two-dimensional axial pretraining, three-dimensional anatomical pretraining and report-guided semantic alignment. This training strategy supports slice-level, volume-level and vision-language analysis. Across five downstream task families (segmentation, classification, registration, vision-language understanding and clinical retrieval), FlexiCT matches or exceeds prior task-specific approaches on multiple benchmarks. Its embeddings further organize CT scans along gradients associated with various tumor stages, suggesting that CT foundation models can capture imaging features relevant to disease phenotype characterization. Code is available at this https URL
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2605.21906 [cs.CV] |
| (or arXiv:2605.21906v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2605.21906 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yuheng Li [view email]
[v1]
Thu, 21 May 2026 02:28:05 UTC (8,829 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.