Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation
Quick Answer
The Crayotter model, utilizing Group-Relative Preference Backpropagation (GRPB), enhances long-horizon video editing by improving subjective feedback handling, surpassing proprietary systems on AgenticVBench.
Quick Take
It effectively transforms ordinal comparisons into actionable insights, leading to better editing outcomes and quality.
Key Points
- GRPB converts subjective video editing feedback into ordinal comparisons for improved decision-making.
- Crayotter model achieves superior performance on AgenticVBench compared to proprietary systems.
- The approach includes a lagged allocator to prevent immediate biases in editing judgments.
- Training involves a project-disjoint, horizon-stratified suite of realistic editing tasks.
- Code and supporting materials are publicly available for further research and development.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Long-horizon video editing agents receive final-product feedback only after many interdependent decisions. Yet editing quality is subjective, admits multiple valid solutions, and is not meaningfully calibrated across heterogeneous requests, making a global scalar objective both ambiguous and temporally uninformative. Our key observation is that fixing the request, materials, and production constraints converts this subjective objective into an ordinal comparison among directly comparable alternatives. We introduce Group-Relative Preference Backpropagation (GRPB), which transforms same-task rankings into zero-sum advantages and redistributes them as bounded credit over semantic editing segments. A lagged allocator and guarded transmission prevent current judgments or unreliable estimates from directly shaping the same rollout group. We manually construct a project-disjoint, horizon-stratified suite of realistic editing tasks for training and controlled evaluation. Across matched baselines, credit interventions, external benchmarking, and blinded human evaluation, GRPB improves both editing behavior and rendered products. The resulting 9B Crayotter model surpasses several proprietary systems on AgenticVBench, supporting task-local preference reduction as a practical approach to learning from subjective, delayed outcomes. Code and all supporting materials are publicly available at this https URL.
| Comments: | 12 pages, 3 figures |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2608.02694 [cs.CL] |
| (or arXiv:2608.02694v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2608.02694 arXiv-issued DOI via DataCite |
Submission history
From: Lecheng Yan [view email]
[v1]
Mon, 3 Aug 2026 10:41:25 UTC (489 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.