MentalThink: Shaping Thoughts in Mental SVG World
Quick Answer
MentalThink introduces a visual-symbolic reasoning paradigm for Multimodal LLMs, enabling them to generate and interpret SVG code for enhanced spatial reasoning.
Quick Take
The model achieves notable benchmarks, scoring 55.1% on VSIBench and 76.0% on MindCube, showcasing its capacity for dynamic visual reflection and scene construction.
Key Points
- MentalThink utilizes a think-with-SVG pipeline for multi-turn reasoning.
- The model combines Supervised Fine-Tuning and Reinforcement Learning for training.
- It externalizes spatial hypotheses through structured vector sketches.
- Extensive evaluations show superior performance on spatial understanding benchmarks.
- The approach mimics human mental imagery processes effectively.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Kangheng Lin, Jisheng Yin, Dingming Li, En Yu, Yana Wei, Han Zhou, Liang Zhao, Hongyu Zhou, Hongbo Peng, Jianjian Sun, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Jingyu Wang
Abstract:We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink is a think-with-SVG pipeline, where the model learns to generate, render, and interpret scalable vector graphics (SVG) code as an intermediate visual representation for multi-turn reasoning. By creating structured vector sketches, the model can externalize spatial hypotheses, inspect them through deterministic rendering, and reason within a constrained geometric space, effectively mimicking the human process of mental imagery. We instantiate this paradigm through a two-stage training framework, combining Supervised Fine-Tuning (SFT) for SVG syntactic alignment with multi-turn Reinforcement Learning (RL) to encourage iterative inspection, revision, and refinement of intermediate visual hypotheses. Extensive evaluations demonstrate that MentalThink achieves superior performance on spatial understanding and reasoning benchmarks (e.g., 55.1% on VSIBench, 76.0% on MindCube), showing that executable vector graphics provide a verifiable visual workspace for dynamic perspective taking, visual reflection, and compositional scene construction.
| Comments: | 17 pages, 6 figures |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.03530 [cs.AI] |
| (or arXiv:2607.03530v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.03530 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kangheng Lin [view email]
[v1]
Fri, 3 Jul 2026 17:59:58 UTC (3,367 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.