Enabling Fully Integer-Only Inference for Lightweight Detection Transformers
Quick Answer
I-LW-DETR is the first fully integer-only lightweight DETR, achieving efficient inference on NPUs and microcontrollers.
Quick Take
It reduces model size by approximately 3.6x and computational cost by over an order of magnitude, with only moderate accuracy degradation. This addresses the deployment challenges of Vision Transformer detectors in resource-constrained environments.
Key Points
- I-LW-DETR executes all operations in integer arithmetic, including nonlinearities.
- Introduces scale-preserving split convolution and SD-ShiftGELU for improved performance.
- Achieves a model size reduction of approximately 3.6x.
- Computational cost is reduced by more than one order of magnitude.
- Moderate accuracy degradation is observed across different model scales.
DeepSignal Analysis
What happened
I-LW-DETR is introduced as the first fully integer-only lightweight DETR, designed for efficient deployment on NPUs and microcontrollers. It achieves a model size reduction of approximately 3.6 times and a computational cost decrease by over an order of magnitude, with only moderate accuracy loss.
Key evidence
- I-LW-DETR executes every operation in the forward pass using integer arithmetic, addressing compatibility issues with existing quantized detectors.
- The model incorporates a scale-preserving split convolution, a sign-dependent GELU approximation, and a constrained Shiftmax for stable Softmax normalization.
- Experimental results indicate that the quantization pipeline consistently produces efficient integer-only models across various scales, with moderate accuracy degradation.
Why it matters
The development of I-LW-DETR is significant as it enables the deployment of Vision Transformer detectors in resource-constrained environments, where traditional models struggle due to their reliance on non-integer operations. This advancement could enhance the applicability of advanced detection models in edge computing and IoT devices, where computational resources are limited.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Vision Transformer detectors now approach the accuracy of CNNs but remain difficult to deploy on NPUs and microcontrollers because key components, including deformable attention, feature fusion, and nonlinear activation functions, are not natively compatible with integer arithmetic. Existing quantized detectors either retain operators such as Softmax, GELU, and LayerNorm or focus on heavyweight backbones, leaving lightweight detection transformers without an end-to-end integer implementation. We address this gap with I-LW-DETR, the first fully integer-only lightweight DETR, in which every operation in the forward pass, including transformer nonlinearities, is executed in integer arithmetic. I-LW-DETR is built upon three key components: a scale-preserving split convolution that assigns independent activation scale to each branch of the multi-scale projector; SD-ShiftGELU, a sign-dependent GELU approximation that preserves element-wise behavior while avoiding the accuracy degradation; and a constrained Shiftmax that maintains stable Softmax normalization. Experimental results demonstrate that the proposed quantization pipeline consistently produces efficient fully integer-only models across different model scales. Across all model scales, the proposed pipeline incurs only a moderate accuracy degradation while reducing the model size by approximately $3.6\times$ and the computational cost by more than one order of magnitude.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.24981 [cs.CV] |
| (or arXiv:2607.24981v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.24981 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Martyna Poreba [view email]
[v1]
Mon, 27 Jul 2026 18:33:18 UTC (6,254 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.