A Shared Latent for Partially-Labeled Multi-Task Facial Affect Recognition
Quick Answer
The study proposes a novel approach to partially-labeled multi-task facial affect recognition by utilizing a shared latent variable, improving expression macro-F1 from 0.403 to 0.446 on s-Aff-Wild2.
Quick Take
This method effectively leverages cross-task signals, achieving a combined multi-task score of 1.679 on validation, while addressing representational failures in rare classes.
Key Points
- Introduces marginalization over a shared affect latent for multi-task learning.
- Improves expression macro-F1 score from 0.403 to 0.446 using a single backbone.
- Achieves a combined multi-task score of 1.679 on the validation set.
- Addresses representational failures in rare-class predictions, not loss shaping.
- Utilizes s-Aff-Wild2 dataset with only 37% of frames fully labeled.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Facial affect in the wild is naturally multi-task: valence-arousal, discrete expressions, and facial action units describe the same face. Yet real corpora annotate these tasks only partially and unevenly, so most systems mask the missing labels or impute pseudo-labels and forgo the cross-task signal. We instead cast partially-labeled multi-task learning as marginalization over a shared affect latent: one variational bottleneck mediates all three task decoders, so a frame annotated for one task shapes the representation the others use, and the masked objective reappears as the reconstruction term of an evidence lower bound. On s-Aff-Wild2, where only 37% of frames carry all three labels, the classes are severely imbalanced, and pretraining on the source data is disallowed, we isolate where this coupling acts. On a single backbone it lifts expression macro-F1 from 0.403 for a dedicated specialist to 0.446, which the masked-loss model does not reach; a second, near-peer backbone with decorrelated errors then breaks an action-unit ceiling that external action-unit data could not, while valence-arousal stays within noise. Every gain is disciplined by a matched-control negative; together these controls indicate that the rare-class failure is representational, not a matter of loss shaping. As each task's source is chosen on the evaluation split, we report the assembled result, a combined multi-task score of 1.679 on validation, as an in-sample endpoint and rest our conclusions on the controlled comparisons; a small, regime-dependent transfer of the expression advantage to AffectNet and RAF-DB is presented as exploratory rather than conclusive.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.16285 [cs.CV] |
| (or arXiv:2607.16285v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16285 arXiv-issued DOI via DataCite |
Submission history
From: Van Thong Huynh [view email]
[v1]
Sat, 11 Jul 2026 04:55:12 UTC (1,087 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.