MiLSD: A Micro Line-Segment Detector for Resource-Constrained Devices
Quick Answer
MiLSD is a micro line-segment detector designed for resource-constrained devices, achieving a significant accuracy improvement on the ShanghaiTech Wireframe benchmark from 10.6 to 24.1 sAP10 with a one-megabyte activation budget.
Quick Take
The model leverages 8-bit quantization to maintain performance while optimizing for low-cost MCUs, making it suitable for embedded vision systems.
Key Points
- MiLSD achieves a significant accuracy boost on ShanghaiTech Wireframe with 1 MB budget.
- 8-bit quantization preserves performance, while 4-bit leads to degradation in angle regression.
- The model systematically compares three output representations within a compact backbone.
- Quantization-aware training recovers only part of the performance loss in 4-bit quantization.
- MiLSD is tailored for MCU-level constraints, enhancing embedded vision applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Line segment detection is a key building block in visual SLAM, 3D reconstruction, and industrial inspection. Recent deep learning methods have greatly improved accuracy, yet even the smallest models require several megabytes of memory, exceeding low-cost MCU capacity. This work investigates the maximum achievable accuracy under a sub-megabyte budget. We propose MiLSD, a detector tailored for MCU-level constraints, and systematically compare three output representations within a compact fully-convolutional backbone. Our study shows that the proposed F-Clip center-with-length-and-angle formulation learns most effectively at small model sizes. We find that 8-bit quantization preserves full-precision performance, while 4-bit quantization causes significant degradation, particularly in angle regression, with quantization-aware training recovering only part of the loss. With a one-megabyte activation budget and inference enhancements including sub-pixel decoding, test-time augmentation, and a lightweight verifier, MiLSD improves sAP10 on ShanghaiTech Wireframe from 10.6 (25k parameters, 0.25 MB) to 24.1 within 1 MB. Rather than competing with GPU-scale parsers, we map the accuracy memory trade-off across representations, bit-widths, capacities, and post-processing strategies for embedded vision systems.
| Comments: | 10 pages, 12 figures, 5 tables |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2607.06600 [cs.CV] |
| (or arXiv:2607.06600v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.06600 arXiv-issued DOI via DataCite |
Submission history
From: Parsa Hassani Shariat Panahi [view email]
[v1]
Tue, 7 Jul 2026 00:04:50 UTC (1,421 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.