VCR-Bench: A Modular Open-Source Benchmark for Video Classification Robustness
Quick Answer
VCR-Bench is a new open-source benchmark for video classification that standardizes evaluation across 30 models and 14 adversarial attacks.
Quick Take
It facilitates reproducibility in robustness studies by providing a unified framework for video loading, metrics, and logging. The benchmark has been tested on the Kinetics-400 dataset, reporting key performance metrics such as accuracy and attack success rates.
Key Points
- Integrates 30 video classification models and 14 adversarial attacks.
- Standardizes video loading and evaluation metrics for reproducibility.
- Evaluated on Kinetics-400, reporting clean accuracy and attack success rates.
- Includes documented installation and component-extension interfaces.
- Facilitates research in video classification robustness.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Robustness of image classification has several benchmarks, but their video counterparts are absent. In video classification temporal dimension introduces additional degrees of freedom for adversarial attacks, defenses, and preprocessing. Temporal sampling, perturbation budgets, and metric aggregation also interact in ways with no direct analogue in the image setting. Therefore, robustness for video classifiers is studied across scattered, incompatible implementations, making reported numbers hard to reproduce and analyze. We introduce VCR-Bench, a modular open-source benchmark framework that standardizes video loading, wrappers for classifiers, adversarial attacks and defenses, perceptual metrics, configuration presets, and result logging. VCR-Bench currently integrates 30 video classification models, 14 adversarial attacks, and 10 defense wrappers under a common evaluation protocol. We evaluate representative video classifiers, attacks, and defenses on Kinetics-400 subset, reporting clean accuracy, attack success rate, perceptual quality, runtime, and memory usage. VCR-Bench is released with documented installation, reproducible run presets, component-extension interfaces, and scripts for reproducing the reported results at this https URL.
| Comments: | 6 pages,1 figure, accepted at ACM MM 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.08936 [cs.CV] |
| (or arXiv:2610.08936v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08936 arXiv-issued DOI via DataCite (pending registration) |
|
| Related DOI: | https://doi.org/10.1145/3767308.3834753
DOI(s) linking to related resources |
Submission history
From: Maksim Plinskiy [view email]
[v1]
Tue, 6 Oct 2026 18:03:58 UTC (1,292 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.