Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge
Quick Answer
This paper shows that TalTech's Voxtral Mini and Small models excelled in the Beyond Transcription Challenge, generating SOAP notes from doctor-patient conversations without transcription.
Quick Take
They achieved first place in both tracks, demonstrating low hallucination rates and improved robustness through fine-tuning and reinforcement learning.
Key Points
- Voxtral Mini and Small adapted with LoRA fine-tuning and DAPO reinforcement learning.
- Ranked first in both lightweight and heavyweight tracks of BeTraC.
- Achieved the lowest hallucination rate among all submissions.
- Fine-tuning on text transcripts improved robustness for speech input.
- Utilized Open Medical Concept F1 as a challenge metric.
Paper Resources
📖 Reader Mode
~2 min readAbstract:This paper describes TalTech's submissions to the Beyond Transcription Challenge (BeTraC), which requires generating SOAP notes directly from long doctor-patient conversation recordings, without intermediate transcription. After screening open-weight speech LLMs for long-audio robustness, we adapted Voxtral Mini (lightweight track) and Voxtral Small (heavyweight track) with LoRA supervised fine-tuning followed by DAPO reinforcement learning that uses the challenge metric, Open Medical Concept F1, as its reward. Our systems ranked first in both tracks, and an independent LLM-as-a-judge evaluation showed the lowest hallucination rate among all submissions, indicating that reinforcement learning against a concept-matching metric need not compromise factual reliability. We also find that fine-tuning on text transcripts transfers well to speech input and appears to improve robustness on out-of-domain real recordings.
| Comments: | SLT 2026 BeTraC |
| Subjects: | Computation and Language (cs.CL); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2607.17230 [cs.CL] |
| (or arXiv:2607.17230v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.17230 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tanel Alumäe [view email]
[v1]
Sun, 19 Jul 2026 12:49:35 UTC (81 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.