CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
Quick Answer
CLIP-CC-Bench introduces a novel evaluation suite for long-form video descriptions, utilizing 90-second clips from 5 hours of movie content paired with expert-written references.
Quick Take
It employs five -based models for reliable semantic matching, assessing 17 state-of-the-art video-language models and providing standardized evaluation tools for reproducibility.
Key Points
- CLIP-CC-Bench evaluates paragraph-level video descriptions, addressing gaps in existing benchmarks.
- The suite is based on 5 hours of movie content segmented into 90-second clips.
- It employs coarse and fine-grained semantic matching methodologies for evaluation.
- 17 state-of-the-art video-language models were assessed, with Borda-aggregated rankings reported.
- Standardized evaluation scripts and tools are released to enhance reproducibility.
DeepSignal Analysis
What happened
CLIP-CC-Bench has been introduced as a new evaluation suite specifically designed for long-form video descriptions. It utilizes 90-second clips from a total of 5 hours of movie content, each paired with expert-written paragraph references. The suite evaluates 17 video-language models using advanced LLM-based embedding models to ensure reliable semantic matching.
Key evidence
- The evaluation suite is built from 5 hours of movie content segmented into 90-second clips, each with an expert-written reference.
- CLIP-CC-Bench employs five state-of-the-art LLM-based models to enhance reliability and reduce bias in semantic matching.
- The framework evaluates 17 video-language models, reporting their rankings and average scores on CLIP-CC-Bench.
Why it matters
This evaluation suite addresses a significant gap in the benchmarking of video-language models, which has previously focused on short clips and single-sentence metrics. By providing a structured method for assessing long-form descriptions, CLIP-CC-Bench aims to improve the understanding of model performance in generating coherent and contextually relevant video descriptions. This advancement could lead to better applications in multimedia content analysis and retrieval.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions. We introduce CLIP-CC-Bench, an evaluation suite for long-form video description built from 5 hours of movie content segmented into 90-second clips, each paired with an expert-written paragraph-style reference. The evaluation suite employs an ensemble of five state-of-the-art LLM-based embedding models to increase reliability and mitigate single-model bias, and applies two complementary methodologies: (i) coarse-grained semantic matching and (ii) fine-grained semantic matching to compare model-generated descriptions against CLIP-CC-Bench references. Using this framework, we evaluate 17 state-of-the-art video-language models and report both their Borda-aggregated rankings and their average scores on CLIP-CC-Bench. We further quantify the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability. We release standardized evaluation scripts, model outputs, and aggregation tools at this https URL to support reproducibility. CLIP-CC-Bench provides a practical evaluation framework for long-form video description, filling a gap left by existing short-clip and QA-only benchmarks.
| Comments: | Accepted and presented at EvalMG 2026, the Second Workshop on Evaluation for Multimodal Generation, co-located with ACM SIGIR 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR); Multimedia (cs.MM) |
| Cite as: | arXiv:2608.04302 [cs.CV] |
| (or arXiv:2608.04302v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.04302 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chulwoo Pack [view email]
[v1]
Wed, 5 Aug 2026 00:20:23 UTC (908 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.