JIVEAdapter: A Multi-Task Additive Low-Rank Adapter via Joint and Individual Variation Explained (JIVE)
Quick Answer
JIVEAdapter introduces a multi-task additive low-rank adapter that separates shared and task-specific signals, achieving competitive performance on GLUE and SuperGLUE benchmarks with DeBERTaV3-base.
Quick Take
It allows for efficient parameter reuse across tasks without retraining the shared components, optimizing both interpretability and adaptability in model fine-tuning.
Key Points
- JIVEAdapter decomposes weight updates into Joint and Individual structures for better interpretability.
- It penalizes Individual structures to maintain near-orthogonality with the Joint structure.
- Competitive performance on GLUE and SuperGLUE benchmarks without extra modules like MoE.
- Allows for adaptive rank allocation across shared and per-task pools.
- Frozen Joint structure can be reused for new tasks, reducing retraining costs.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Parameter-efficient fine-tuning adapts pretrained models at a fraction of the cost of full fine-tuning, yet most low-rank adapters are single-task and represent each weight update multiplicatively, leaving no explicit account of what is shared across tasks and what is task-specific. We introduce JIVEAdapter, a multi-task "additive" low-rank adapter inspired by statistical Joint and Individual Variation Explained (JIVE). JIVEAdapter decomposes every weight update into a Joint structure shared across all tasks plus a per-task Individual structure, penalizes the Individual structures to be near-orthogonal to the Joint so shared and task-specific signal stay "interpretable" and separated, and allocates rank adaptively across a shared Joint pool and a per-task Individual pool. The Joint is learned once, jointly over a task group or incrementally, one task at a time, then frozen and reused as a prior for new tasks without retraining the shared part. On GLUE and SuperGLUE with DeBERTaV3-base, JIVEAdapter is competitive with strong single-task and multi-task low-rank baselines at a matched per-task effective rank, without extra modules such as MoE, and when a related held-in task exists its frozen Joint serves a held-out task by reusing that task's Individual with only a cheap per-direction scale, otherwise training a small new one.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07036 [cs.AI] |
| (or arXiv:2610.07036v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07036 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sara Abdali [view email]
[v1]
Sun, 4 Oct 2026 21:09:25 UTC (290 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.