Large Language Models as Unified Multimodal Learners for Clinical Prediction
Quick Answer
This study presents a unified approach for clinical prediction by converting diverse patient data into a single natural language sequence, leveraging pretrained models like Llama 3.1 and Gemma.
Quick Take
The method outperforms traditional multimodal systems and a clinical gradient boosting model in predicting graft failure, simplifying system architecture while maintaining high accuracy across three distinct tasks.
Key Points
- Unified textual serialization matches or exceeds task-specific multimodal baselines.
- Outperformed clinical gradient boosting model in graft failure prediction.
- Evaluated across in-hospital mortality, graft failure, and emergency triage tasks.
- No architectural modifications needed for fusion in the proposed method.
- Simplifies clinical prediction systems significantly while maintaining performance.
DeepSignal Analysis
What happened
The study introduces a method for clinical prediction that transforms various patient data into a single natural language sequence. This approach utilizes pretrained models like Llama 3.1 and Gemma, showing improved performance over traditional multimodal systems and a clinical gradient boosting model in predicting graft failure.
Key evidence
- The proposed method converts diverse patient data into a single natural language sequence, simplifying the architecture of clinical prediction systems.
- The evaluation includes three distinct tasks: in-hospital mortality, graft failure prediction, and emergency triage classification, demonstrating the method's versatility.
- Results indicate that the unified textual serialization approach matches or exceeds the performance of task-specific multimodal baselines and outperforms a clinical gradient boosting model for graft failure.
Why it matters
This research highlights a shift towards a more streamlined approach in clinical prediction, potentially reducing the complexity of system architectures. By leveraging a single serialization method, it may facilitate easier implementation across various clinical settings, which could enhance predictive accuracy and efficiency in patient management.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities. Yet most clinical prediction systems still rely on task-specific fusion architectures, pairing dedicated encoders for each modality with learned combination mechanisms that must be re-engineered for every new task and clinical setting. We propose a simpler alternative: convert all patient data, regardless of modality, into a single natural language sequence and fine-tune a pretrained language model end-to-end, with no architectural modification for fusion. We evaluate this approach across three clinically distinct prediction tasks: in-hospital mortality on MIMIC-III, graft failure prediction using longitudinal data from a German transplant center, and emergency triage classification from ambulance records - comparing encoder-based (ModernBERT) and decoder-based (Llama 3.1, Gemma, DeepSeek-R1-Qwen, Qwen3) fine-tuning against established multimodal baselines and, for graft failure, a gradient boosting model currently used in clinical practice for post-transplant patient management. Across all three tasks, unified textual serialization matches or exceeds task-specific multimodal baselines, and outperforms the clinically deployed gradient boosting system on graft failure prediction. These results indicate that a single serialization-based paradigm, without bespoke fusion architectures, is sufficient for multimodal clinical prediction - substantially reducing system complexity while matching or exceeding specialized designs.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.15380 [cs.CL] |
| (or arXiv:2607.15380v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15380 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Roland Roller [view email]
[v1]
Thu, 16 Jul 2026 18:28:23 UTC (50 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.