Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine Translation
Quick Answer
This study evaluates local LLMs like llama3.2:3b and mistral:latest for machine translation across nine EU languages, revealing that dedicated MT systems outperform LLMs, particularly for Germanic languages.
Quick Take
Few-shot prompting benefits some models but not others, while family-scope prompts expose weaknesses in smaller models.
Key Points
- Evaluated English-to-Romance and English-to-Germanic translations using FLORES devtest.
- Compared three local against OPUS-MT and NLLB-200 baselines.
- Few-shot prompting improved performance for mistral:latest and qwen2.5:14b.
- Family-scope prompting is feasible for stronger LLMs but reveals issues in smaller models.
- Results emphasize the need for diverse evaluation metrics in LLM translation.
DeepSignal Analysis
What happened
This study assesses local large language models (LLMs) for machine translation across nine EU languages, comparing them to dedicated machine translation (MT) systems. It evaluates various prompting strategies, including few-shot and family-scope prompts, revealing that dedicated MT systems generally outperform LLMs, particularly for Germanic languages.
Key evidence
- The study evaluates English-to-Romance and English-to-Germanic translation using the FLORES devtest split for nine EU languages.
- Three local instruction-tuned LLMs, llama3.2:3b, mistral:latest, and qwen2.5:14b, are compared against dedicated MT systems from OPUS-MT and NLLB-200.
- Results indicate that few-shot prompting benefits some models like mistral:latest but negatively impacts llama3.2:3b, while family-scope prompts expose weaknesses in smaller models.
Why it matters
Understanding the performance of local LLMs in machine translation is crucial as these models are increasingly used in practical applications. The findings highlight the limitations of LLMs compared to dedicated MT systems, particularly in specific language families. This research also emphasizes the importance of evaluating translation models based on prompt strategies and their effectiveness across different languages.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models (LLMs) are increasingly used as general-purpose translation systems, but their behavior is usually evaluated under a single prompt shape: translate one source sentence into one target language. In practice, users may ask for one target language, for several related languages at once, or for translations conditioned on examples. This paper studies prompt scope and demonstration selection as experimental variables for local LLM machine translation. We evaluate English-to-Romance and English-to-Germanic translation on the full FLORES devtest split for nine official European Union languages. We compare three local instruction-tuned LLMs, llama3.2:3b, mistral:latest, and qwen2.5:14b, against dedicated MT baselines from OPUS-MT and NLLB-200. We test zero-shot prompting and k=5 few-shot prompting with random, lexical-similarity, and embedding-similarity demonstration selection. We also compare single-target prompts with JSON-formatted family-scope prompts that request all languages in a family at once. Results show that dedicated MT systems remain strongest overall, especially for Germanic languages. Few-shot prompting helps mistral:latest and qwen2.5:14b, but hurts llama3.2:3b; embedding retrieval is best on average for the stronger LLMs, but its advantage over random and lexical examples is modest. Family-scope prompting is feasible for stronger local LLMs but exposes structured-output failures in smaller models. These findings motivate evaluating LLM translation not only by language pair and metric, but also by prompt scope, retrieval strategy, and multi-target compliance.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.26286 [cs.CL] |
| (or arXiv:2607.26286v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.26286 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mihael Arcan [view email]
[v1]
Tue, 28 Jul 2026 21:26:36 UTC (18 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.