LegoQ: Density-Matrix Representation Learning with Spectral-Spatial State Transitions for Hyperspectral Classification
Quick Answer
The paper introduces LegoQ, a density-matrix representation learning framework for hyperspectral image classification, achieving 96.20% accuracy on Indian Pines and 97.52% on WHU-Hi-LongKou.
Quick Take
It effectively addresses mixed pixels and spectral ambiguity without requiring quantum hardware, offering improved diagnostics through sample-level metrics.
Key Points
- LegoQ uses a classical density-matrix framework for hyperspectral image classification.
- Achieved 96.20% accuracy on Indian Pines and 97.52% on WHU-Hi-LongKou.
- Aggregates group states and compares them with learnable class-prototype density matrices.
- Offers improved diagnostics like von Neumann entropy and prototype fidelity.
- Provides a practical alternative to vector-only hyperspectral classification.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Hyperspectral image classification is complicated by mixed pixels, spectral ambiguity, class imbalance, and limited annotations. Most current classifiers encode a pixel or patch as a deterministic vector and apply a linear or multilayer softmax head. Although effective for discrimination, this representation does not directly expose how mixed or uncertain a sample is. This paper presents \method, a classical density-matrix representation learning framework for hyperspectral images. The spectral bands are divided into groups and each group is mapped to a positive semi-definite, Hermitian, trace-normalized matrix state. A composable stack of spectral, spatial, and inter-group transitions then updates the states while repeatedly projecting them back to the valid state set. Instead of flattening the final features, \method\ aggregates the group states and compares them with learnable class-prototype density matrices through Uhlmann fidelity. The normalized eigenspectrum, von Neumann entropy, purity, and prototype fidelity provide sample-level diagnostics that are unavailable from a conventional vector head. On Indian Pines, ten runs yield an overall accuracy of $96.20\pm0.70\%$, an average accuracy of $95.57\pm1.29\%$, and a kappa coefficient of $95.66\pm0.80\%$. On WHU-Hi-LongKou, the best of ten runs reaches $97.52\%$ overall accuracy. Classification maps and feature projections show that the transition stack produces compact and better separated class structures. The results support constrained matrix-state learning as a practical alternative to vector-only hyperspectral classification without requiring quantum hardware.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.28970 [cs.CV] |
| (or arXiv:2607.28970v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.28970 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Weijia Cao [view email]
[v1]
Fri, 31 Jul 2026 02:47:16 UTC (625 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.