MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models
Quick Answer
MIITA introduces a novel framework for continual learning in small language models, addressing catastrophic forgetting by using memory-based adaptation without updating model parameters.
Quick Take
It effectively utilizes compact prototypes for knowledge retention, improving performance while mitigating forgetting under fixed memory constraints.
Key Points
- MIITA utilizes compact correction-direction prototypes for knowledge retention.
- The framework retrieves experiences at inference time using semantic cues.
- It enables non-destructive reuse of past supervision without model updates.
- Extensive experiments show consistent performance improvements in supervised CL.
- The approach mitigates forgetting within fixed memory budgets.
DeepSignal Analysis
What happened
The MIITA framework was proposed to enhance continual learning in small language models by addressing catastrophic forgetting. It employs memory-based adaptation techniques that do not require updating model parameters, utilizing compact prototypes for knowledge retention. This approach aims to improve performance while adhering to fixed memory constraints.
Key evidence
- MIITA stores supervised experiences as compact correction-direction prototypes, which are retrieved at inference time using semantic and uncertainty-based cues.
- The framework allows for non-destructive reuse of past supervision without necessitating updates to the model's backbone or prompt extensions.
- Extensive experiments demonstrate that MIITA consistently enhances final performance and reduces forgetting within fixed memory budgets.
Why it matters
Continual learning is crucial for small language models to adapt to changing real-world requirements, especially in environments with limited resources. Traditional methods often lead to catastrophic forgetting when model parameters are updated. MIITA's innovative approach could provide a viable solution, allowing these models to retain knowledge effectively while maintaining performance, which is essential for practical applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Continual learning (CL) is essential for small language models (SLMs) to adapt to evolving real-world needs in resource-constrained deployments. However, directly updating their limited parameter space causes catastrophic forgetting. While memory-based methods naturally address this by decoupling knowledge retention from parameters, existing approaches designed for large language models (LLMs) rely on abundant storage and strong in-context reasoning that SLMs lack. To address these challenges, we propose MIITA, a Memory-Induced Inference-Time Adaptation framework for supervised CL under constrained storage. MIITA stores supervised experiences as compact correction-direction prototypes with semantic anchors, and retrieves them at inference time using semantic and uncertainty-based cues. The retrieved directions are applied through gated temporary hidden-state adaptation, enabling non-destructive reuse of past supervision without backbone updates, prompt extensions, or test-time backpropagation. A local theoretical analysis links this design to first-order loss reduction, uncertainty-guided retrieval, and directional coverage for retaining old-stage knowledge. Extensive experiments across diverse supervised CL settings show that MIITA consistently improves final performance and mitigates forgetting under fixed memory budgets.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.22556 [cs.AI] |
| (or arXiv:2607.22556v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22556 arXiv-issued DOI via DataCite |
Submission history
From: Dong Li [view email]
[v1]
Wed, 20 May 2026 03:03:04 UTC (148 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.