Group-wise Supervision with Focal-Dice Loss for Long-Tailed Indoor Semantic Occupancy Prediction
Quick Answer
The proposed Group-UFD Occ method enhances long-tailed indoor semantic occupancy prediction by 11.38% over baseline using hierarchical supervision and Unified Focal-Dice loss, effectively addressing the challenges posed by diverse object categories.
Quick Take
This approach employs fine-grained semantic grouping and multi-scale prediction heads to improve tail-class feature learning.
Key Points
- Introduces Group-UFD Occ for improved indoor semantic occupancy prediction.
- Achieves 11.38% relative improvement on the EmbodiedScan dataset.
- Utilizes hierarchical semantic supervision and synergistic loss optimization.
- Employs fine-grained semantic grouping and multi-scale prediction heads.
- Focuses on hard samples with Unified Focal-Dice loss at the voxel level.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Recently, 3D semantic occupancy prediction has garnered increasing attention for understanding the indoor scene. However, unlike structured outdoor environments, indoor scenes feature a high diversity of object categories that exhibit a severe long-tailed distribution, which has become a core bottleneck limiting the performance of existing models. To tackle this challenge, we propose a novel method, Group-UFD Occ, based on hierarchical semantic supervision and synergistic loss optimization. At the architectural level, we introduce a fine-grained semantic grouping strategy and design multi-scale, parallel ``main-expert'' prediction heads to guide the model in efficiently learning tail-class features through deep regularization. At the optimization level, we introduce the Unified Focal-Dice (UFD) loss. This synergistic loss function dynamically focuses on hard samples at the per-voxel level. Meanwhile, it simultaneously optimizes the geometric integrity of predicted objects from a region-based perspective. We conducted experiments on the large-scale EmbodiedScan dataset. The results demonstrate that our method yields a relative improvement of 11.38\% over the baseline, with substantial accuracy gains in several critical long-tailed categories.
| Comments: | 8 pages, 2 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.28935 [cs.CV] |
| (or arXiv:2607.28935v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.28935 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Qi Zheng [view email]
[v1]
Fri, 31 Jul 2026 01:32:45 UTC (1,657 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.