The JEPA Predictor: A Transferable Operator for Occluded Feature Completion
Quick Answer
The JEPA Predictor enhances occluded feature completion by acting as a transferable operator across various encoder families, achieving significant accuracy improvements on ImageNet-9 and Stanford Dogs.
Quick Take
By integrating frozen predictors from I-JEPA and V-JEPA with non-JEPA models like CLIP, accuracy on heavily occluded images increased from 15.9% to 52.1%, demonstrating the effectiveness of this approach without requiring retraining.
Key Points
- JEPA Predictor closes accuracy gaps in occluded feature completion across encoder families.
- Accuracy on ImageNet-9 improved significantly with CLIP and I-JEPA integration.
- Stanford Dogs classification accuracy increased from 15.9% to 52.1% using frozen predictors.
- The method requires no retraining, fitting linear probes per mask fraction.
- Performance benefits grow with higher mask fractions in occluded scenarios.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Joint-Embedding Predictive Architectures (JEPAs) train a predictor jointly with their encoder, but downstream deployment discards the predictor and reads features from the encoder alone. The predictor is, by construction, a learned operator from visible-context features to features at masked positions, the structure a partial-view classifier needs. We show that this operator is portable across encoder families. We first establish that, at heavy mask, retaining the frozen predictor on a JEPA encoder substantially closes the accuracy gap against the strongest non-JEPA discriminative baselines. We then bolt the frozen predictors of I-JEPA and V-JEPA 2 onto four non-JEPA hosts (CLIP, DINOv3, DINOv2, MAE) through a single linear projection between feature spaces, fit in closed form on 500 ImageNet-1k images. Across both ImageNet-9 and Stanford Dogs and across three mask fractions, the lift over each host's masked-encoder baseline grows monotonically with the mask fraction K in every host-donor pair. CLIP paired with the I-JEPA predictor recovers most of the accuracy that masking removed on ImageNet-9 at heavy occlusion, and lifts fine-grained Stanford Dogs from 15.9% to 52.1% (+36 pp). The mechanism is identifiable: the projection pays a fixed cost on visible patches and the predictor provides a growing benefit on masked patches; the benefit dominates the heavy-occlusion regime. At low K on fine-grained classification the projection cost exceeds the benefit, defining the boundary where the linear bridge breaks down. The frozen JEPA predictor functions as a portable operator for occluded feature completion across encoder families, requiring no retraining of either model while fitting matched linear probes per mask fraction.
| Comments: | 12 pages, 3 figures, 5 tables |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.16274 [cs.CV] |
| (or arXiv:2607.16274v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16274 arXiv-issued DOI via DataCite |
Submission history
From: William Nguyen [view email]
[v1]
Wed, 8 Jul 2026 10:20:40 UTC (96 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.