GEB-Bench: Abstract Structures Told in Many Voices
Quick Answer
GEB-Bench introduces a benchmark for evaluating models on abstract structural motifs, revealing a significant gap in cross-voice mapping.
Quick Take
Twelve models were tested, showing that while they excel in recognizing structures within a single voice, they struggle to transfer this understanding across different representations, with errors aligning more with formal geometries than perceptual ones.
Key Points
- GEB-Bench focuses on abstract motifs like self-reference and Mobius twists.
- Models perform better in recognizing structures than in cross-voice mapping.
- Errors correlate more with designed geometries than with perceptual ones.
- Frontier models from different vendors yield similar incorrect answers.
- GEB-Bench is fully generative and includes its evaluation pipeline.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstract structural motif--self-reference, a strange loop, a Mobius twist--in the spirit of Godel, Escher, Bach. Each motif is told in several voices: a natural scene whose composition is the structure, a folk story whose telling enacts it through a mechanically checkable form device, a mathematical theorem, and a programmatic skeleton; surface parameters are declared nuisance variables and never scored. Motifs, voices, and the structural changes between them form a small cross-modal category, and GEB-Bench's tasks are its questions. Evaluating twelve open and proprietary models, we find that abstraction failure is lawful. The central finding is a gap between recognition and cross-voice mapping: models identify a structure within one voice far better than they carry it across voices; every model pays this tax, and mapping strong enough to narrow it appears only at the frontier tier. Two patterns support it. Errors align more strongly with the designed formal geometry than with measured perceptual geometries, and frontier models from different vendors converge on the same wrong answers; and surface complexity taxes every model that reads structure, with capacity buying headroom rather than immunity. GEB-Bench is fully generative and released with its pipeline.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Logic in Computer Science (cs.LO) |
| Cite as: | arXiv:2608.04111 [cs.CV] |
| (or arXiv:2608.04111v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.04111 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tong Zhang [view email]
[v1]
Tue, 4 Aug 2026 18:05:41 UTC (22,018 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.