Stochastic Meta-Unlearning: Bridging Language Backbone and Multimodal Unlearning
Quick Answer
This paper shows that Stochastic Meta-Unlearning (SMU) enhances unlearning in vision-language models (VLMs) by leveraging VLM-level feedback, achieving a 10.52 point reduction in average Forget accuracy and improvements of 20.10 and 17.01 points in Retain and Test accuracy, respectively.
Quick Take
This method demonstrates better reliability and transferability in unlearning tasks across different datasets and models.
Key Points
- SMU employs a bilevel framework for effective unlearning in .
- Unlearning feedback from VLMs improves the reliability of language backbone updates.
- Experiments show SMU outperforms baselines in forget-retain trade-offs.
- SMU reduces average Forget accuracy by 10.52 points.
- The method is transferable to new forgetting targets and unlearning methods.
DeepSignal Analysis
What happened
Stochastic Meta-Unlearning (SMU) is introduced to improve unlearning in vision-language models (VLMs). It utilizes VLM-level feedback, resulting in a 10.52 point decrease in average Forget accuracy and increases of 20.10 and 17.01 points in Retain and Test accuracy, respectively. SMU shows enhanced reliability and transferability in unlearning tasks across various datasets and models.
Key evidence
- SMU applies a bilevel framework to learn an unlearning-ready initialization, using VLM-level feedback to enhance the unlearning process.
- Experiments indicate that SMU reduces average Forget accuracy by 10.52 points while improving Retain and Test accuracy by 20.10 and 17.01 points, respectively.
- The method demonstrates better transferability to new forgetting targets and different meta-test unlearning methods, suggesting its broader applicability.
Why it matters
The complexity of unlearning in VLMs, which combine language and visual components, necessitates more sophisticated methods. SMU's reliance on VLM-level feedback addresses the inadequacies of text-only feedback, thereby enhancing the reliability of unlearning processes. This advancement could lead to more effective model updates in applications where data privacy and adaptability are critical, such as in personalized AI systems.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Machine unlearning for vision-language models (VLMs) remains underexplored. Unlike language models, VLMs combine a language backbone with visual components, which makes unlearning more complex. There is a surprising phenomenon when moving from single-modality unlearning to VLM unlearning: a target forgotten by the standalone language backbone can still be recovered when image information is given to the full VLM. This shows that text-only feedback is not enough for reliable VLM unlearning. Motivated by this observation, we propose Stochastic Meta-Unlearning (SMU), a bilevel framework that uses VLM-level feedback to learn an unlearning-ready initialization. In the inner loop, SMU applies a few unlearning steps to the language backbone using text data. In the outer loop, SMU recomposes the updated backbone with the frozen VLM and evaluates forgetting and utility at the VLM level. This design makes the unlearning update aware of the final multimodal behavior, while still keeping the update local to the language backbone. Experiments on two VLMs, two multimodal meme datasets, and three baselines show that SMU achieves the best overall forget-retain trade-off. Compared with the strongest baseline for each metric, SMU reduces average Forget accuracy by 10.52 points and improves average Retain and Test accuracy by 20.10 and 17.01 points, respectively. More importantly, SMU also transfers to new forgetting targets and to different meta-test unlearning methods. These results suggest that VLM-level feedback can make language-backbone unlearning more reliable and more transferable for VLMs.
| Comments: | 15 pages |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.18615 [cs.CL] |
| (or arXiv:2607.18615v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.18615 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zijie Liu [view email]
[v1]
Tue, 21 Jul 2026 01:35:33 UTC (955 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.