Strength-Parity Ensembling with Parameter-Isolated Experts for Multi-Task Affect Recognition
Quick Answer
The study introduces a strength-parity ensembling method for multi-task affect recognition, achieving a validation score of 1.7259, significantly surpassing the baseline of 0.45.
Quick Take
By employing parameter-isolated experts, the method maintains decorrelation while enhancing accuracy in valence-arousal estimation and expression recognition tasks. This approach addresses the challenges of ensemble diversity and accuracy in affective computing.
Key Points
- Introduces strength-parity rule for ensemble member selection in affect recognition tasks.
- Achieves a validation score of 1.7259, outperforming the baseline of 0.45.
- Utilizes parameter-isolated experts to maintain decorrelation and enhance accuracy.
- Focuses on joint valence-arousal estimation and expression recognition from single faces.
- Addresses challenges of ensemble diversity under long-tailed label conditions.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Leading entries on the multi-task track of the 11th ABAW challenge rely on heavy ensembling, yet which member is worth adding to an already strong ensemble is rarely made explicit. We study this question for joint valence-arousal estimation, 8-way expression recognition, and 12-way action-unit detection from a single unconstrained face, under partial, long-tailed labels and a rule that forbids pretraining on Aff-Wild2. Building on a shared affect-latent that marginalizes the missing labels across two affect-supervised backbones, we propose a strength-parity rule: an added member lowers the ensemble error only when it is both decorrelated from the current members and a near-peer of them in individual accuracy. The rule exposes a concrete obstacle, as on a single backbone re-seeding and even distinct fine-tuning curricula re-converge to a prediction correlation of 0.98 and add no diversity. Parameter-isolation removes it: confining each adaptation to a disjoint low-rank subspace of a shared backbone yields experts that stay decorrelated at 0.91 while remaining near-peers, the strongest of them an AffectNet-adapted expert. The resulting system raises the overall validation score to 1.6949, against the organizers ConvNeXt-with-MixAugment baseline of 0.45; with per-AU calibration and by pooling the shared-latent heads valence-arousal byproduct as a further near-peer, the strongest configuration reaches 1.7259.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.16290 [cs.CV] |
| (or arXiv:2607.16290v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16290 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Van Thong Huynh [view email]
[v1]
Sun, 12 Jul 2026 17:12:15 UTC (160 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.