Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization
Quick Answer
The proposed multi-objective structured pruning framework optimizes large language models (LLMs) for edge deployment by reducing latency and model size while maintaining performance.
Quick Take
This two-stage method achieves a favorable trade-off, demonstrating improved commonsense reasoning task performance at 37.5% and 50% pruning ratios compared to existing techniques, significantly lowering inference costs.
Key Points
- Introduces a hardware-aware pruning framework targeting latency and model size.
- Employs coarse-grained depth pruning to eliminate entire attention and MLP blocks.
- Utilizes Parallel Bayesian Optimization for optimal layer-wise pruning ratios.
- Achieves better performance on commonsense reasoning tasks than existing methods.
- Reduces inference costs significantly while maintaining model performance.
DeepSignal Analysis
What happened
A new multi-objective structured pruning framework for large language models (LLMs) has been proposed to optimize them for edge deployment. This framework aims to reduce both latency and model size while maintaining performance levels. The method includes a two-stage approach that effectively prunes model components to achieve these goals.
Key evidence
- The proposed framework employs a two-stage method that includes coarse-grained depth pruning and fine-grained layer-wise pruning to optimize latency and model size.
- Experimental results indicate that the method achieves better performance on commonsense reasoning tasks at 37.5% and 50% pruning ratios compared to existing techniques.
- The approach significantly reduces inference costs while maintaining minimal impact on model performance, making it suitable for deployment in resource-constrained environments.
Why it matters
Optimizing LLMs for edge deployment is crucial due to the increasing demand for efficient AI applications in resource-limited settings. By addressing latency and memory constraints, this framework could facilitate broader adoption of LLMs in practical applications. The ability to maintain performance while reducing model complexity is particularly significant for developers and organizations looking to implement AI solutions in real-world scenarios.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However, deploying them in embedded and edge computing environments remains challenging because of strict latency, memory, and energy constraints. Their large parameter counts and computational demands hinder efficient execution on resource-constrained platforms. Although model pruning has emerged as a viable solution for reducing scale while preserving performance, jointly optimizing layers, attention heads, and Multi-Layer Perceptron (MLP) dimensions remains highly complex. Exhaustively exploring this combined design space is computationally expensive and often leads to local optima or unstable configurations. To address these limitations, we propose a hardware-aware, multi-objective structured pruning framework. The proposed two-stage method explicitly targets latency and model size for efficient deployment on edge devices. In the coarse-grained stage, multi-objective depth pruning removes entire attention and MLP blocks to reduce computational load and memory usage. In the subsequent fine-grained stage, Parallel Bayesian Optimization (PBO) searches for the optimal layer-wise pruning ratios for pruning under latency constraints, while importance-based strategies rank the specific components to be pruned within each layer's allocated budget. Experimental results show that our approach reduces model complexity with minimal impact on commonsense reasoning tasks and zero-shot performance. Our method achieves a favorable trade-off among accuracy, latency, and model size, making it suitable for edge deployment. Across multiple LLMs at 37.5% and 50% pruning ratios, the proposed approach achieves better performance on commonsense reasoning tasks than existing methods while significantly reducing inference cost.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.22583 [cs.AI] |
| (or arXiv:2607.22583v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22583 arXiv-issued DOI via DataCite |
Submission history
From: Muhammad Junaid Ali [view email]
[v1]
Mon, 8 Jun 2026 09:38:13 UTC (4,945 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.