NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task
Quick Answer
This paper shows that NAVER LABS re-implements its IWSLT 2025 instruction-following pipeline for the IWSLT 2026 Shared Task using SeamlessM4T-v2-large and Qwen3-4B-Instruct.
Quick Take
The model achieves COMET 0.781 on EN-ZH speech translation and BERTScore-F1 0.346 on the MCIF benchmark, bolstered by 100k synthetic instruction-following examples across ten task types.
Key Points
- Utilizes SeamlessM4T-v2-large as the speech encoder.
- Employs Qwen3-4B-Instruct as the backbone.
- Maintains a three-stage approach from the original design.
- Generates 100k synthetic instruction-following examples for fine-tuning.
- Achieves notable benchmark scores on EN-ZH translation and SQA.
Paper Resources
📖 Reader Mode
~1 min readAbstract:We re-implement the NAVER LABS IWSLT 2025 instruction-following pipeline for the IWSLT 2026 Shared Task (constrained condition, short audio track), adapting it to the mandated components: SeamlessM4T-v2-large as the speech encoder and Qwen3-4B-Instruct as the LLM backbone. The three-stage approach projector alignment, text-only LoRA pre-training, and multimodal merging is preserved from the original design. We additionally construct 100k synthetic instruction-following examples across ten speech-centric task types (10k per task) from the provided corpora, suitable for further Stage 3 fine-tuning. Our primary model achieves COMET 0.781 on EN-ZH speech translation and BERTScore-F1 0.346 on English SQA on the MCIF benchmark.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.05623 [cs.CL] |
| (or arXiv:2607.05623v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05623 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Anand Kamble [view email]
[v1]
Mon, 6 Jul 2026 20:31:41 UTC (82 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.