ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
Quick Answer
ABACUS is a unified vision-language model that excels in object and crowd counting, as well as count-faithful image generation, achieving state-of-the-art results across seven benchmarks without benchmark-specific training.
Quick Take
It incorporates innovations like density-aware adaptive zooming and a cycle-consistent strategy, outperforming both task-specific models and larger generalist models.
Key Points
- ABACUS uses a 3B-parameter unified foundation model for enhanced performance.
- Innovations include density-aware adaptive zooming and boundary-aware count policy.
- Achieves state-of-the-art results across seven benchmarks.
- Outperforms both task-specific specialists and larger generalist models.
- No external annotations required for understanding and generation tasks.
Paper Resources
📖 Reader Mode
~2 min readAbstract:ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generation without any benchmark-specific training required. Our model is built on existing 3B-parameter unified foundation model and is adapted for object localization tasks using three key innovations: density-aware adaptive zooming with objectness maps for spatial grounding; a boundary-aware count policy via GRPO to eliminate crop-boundary errors; and a cycle-consistent GRPO strategy where the understanding branch self-critiques generated outputs, closing the understanding-generation gap without any external annotations. ABACUS achieves state-of-the-art results across seven benchmarks, outperforming both task-specific specialists and larger generalist models.
| Comments: | Under review, webpage: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV) |
| Cite as: | arXiv:2606.23835 [cs.CV] |
| (or arXiv:2606.23835v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2606.23835 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Anindya Mondal [view email]
[v1]
Mon, 22 Jun 2026 18:16:31 UTC (15,107 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.