AnySimLite: A Lightweight Few-Shot Similarity Encoder for On-Device Speech-Adjacent Classification
Quick Answer
AnySimLite is a lightweight similarity encoder designed for on-device speech-adjacent classification, achieving state-of-the-art performance in few-shot settings while using less than 1/250th the model size of the qLLaMA_LoRA-7B baseline.
Quick Take
It effectively combines word-level and character-level channels to minimize memory footprint and maintain low inference latency on edge devices.
Key Points
- AnySimLite combines word-level and character-level channels for enhanced classification.
- Achieves state-of-the-art performance in few-shot settings across multiple tasks.
- Maintains a low memory footprint, using less than 1/250th the size of qLLaMA_LoRA-7B.
- Performance drop remains below 7% even in worst-case scenarios.
- Addresses privacy concerns and inference latency on edge devices like smartphones.
Paper Resources
📖 Reader Mode
~2 min readAbstract:To minimize privacy concerns and inference latency on edge devices like smartphones, lightweight on-device models remain important for end-user applications. Many of these applications involve natural language classification, but deploying multiple specialized models creates a memory footprint challenge. We investigate: Can a single lightweight architecture solve multiple Speech-Adjacent (SA) classification tasks through reduction to a nuanced text similarity formulation? We propose AnySimLite, a lightweight similarity encoder that combines word-level and character-level channels. Together with a dataset transformation strategy, we evaluate AnySimLite across multiple SA classification tasks and show that it consistently achieves state-of-the-art (SOTA) or SOTA-competitive performance in few-shot settings while maintaining a low memory footprint. Even in the worst case, the performance drop remains below 7% while using $<\frac{1}{250}^{\mathrm{th}}$ of the model size of the SOTA qLLaMA_LoRA-7B baseline.
| Comments: | Accepted at Interspeech 2026 |
| Subjects: | Computation and Language (cs.CL); Sound (cs.SD) |
| Cite as: | arXiv:2606.26452 [cs.CL] |
| (or arXiv:2606.26452v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.26452 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sourav Ghosh [view email]
[v1]
Wed, 24 Jun 2026 23:25:28 UTC (483 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.