Real-Time Semantic Segmentation with Optimized RetinaNet Architectures for Embedded Automotive Systems
Quick Answer
This paper shows that The Opt-RetinaSeg architecture optimizes semantic segmentation for embedded automotive systems, achieving 73.9% mIoU at 70.4 FPS on NVIDIA Jetson Xavier NX.
Quick Take
This model offers a 7.4x speedup and 4x size reduction compared to ResNet-50 with minimal accuracy loss, making it suitable for real-time ADAS applications.
Key Points
- Opt-RetinaSeg replaces ResNet-50 with a lightweight feature extractor.
- Achieves 73.9% mIoU and 70.4 FPS on Cityscapes and BDD100K datasets.
- Utilizes structured channel pruning and post-training INT8 quantization.
- 4x reduction in model size with less than 3% accuracy degradation.
- Demonstrates viability for real-time segmentation in automotive perception.
Paper Resources
📖 Reader Mode
~3 min read
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.