Structured Output Collapses Answer Diversity Across 44 Language Models
Quick Answer
The study reveals that requesting JSON format from 44 language models significantly reduces answer diversity, with the modal answer increasing from 41% to 64% in a one-word prompt.
Quick Take
This structured output leads to a more homogeneous response landscape, affecting how models are evaluated and chosen, particularly highlighting the behavior of distinctive models like Claude Fable.
Key Points
- Answer diversity drops from 52 distinct answers to 36 when using JSON.
- Mean answer-choice surprisal decreases from 1.80 to 1.58 bits.
- Six out of 44 models shift toward the modal answer, especially distinctive ones.
- JSON format compresses response significantly compared to plain chat.
- Structured output impacts how language models are consumed in software.
DeepSignal Analysis
What happened
A study analyzed the impact of requesting JSON format from 44 language models on answer diversity. The results indicated that the modal answer increased from 41% to 64% for a one-word prompt, leading to a more uniform response landscape. This shift was particularly pronounced among distinctive models like Claude Fable.
Key evidence
- The modal answer for the unconstrained 'Pick a word' prompt rose from 41% to 64% when responses were requested in JSON format.
- Distinct answers decreased from 52 to 36, and the mean answer-choice surprisal dropped from 1.80 to 1.58 bits.
- Six out of 44 models showed a shift toward the modal answer, with notable changes in models like Claude Fable, which answered 'cerulean' 100% of the time in JSON but 0% in chat.
Why it matters
This finding highlights how structured output formats, such as JSON, can significantly influence the behavior of language models. The increased homogeneity in responses raises questions about the evaluation and selection of models, as it may not reflect their performance in more natural conversational settings. Understanding these dynamics is crucial for developers and researchers aiming to leverage language models effectively.
Paper Resources
📖 Reader Mode
~2 min readAbstract:When a language model must choose one answer from a large space of equally valid options, a format clause -- "Reply with JSON only" -- changes which answer it chooses. We re-run the One-Word Census (arXiv:2607.12796): 31 wide-answer-space category prompts asked of 44 models, now with the reply requested in JSON -- no schema enforcement, no constrained decoding, only the request. Convergence deepens sharply: on the unconstrained "Pick a word" prompt the modal answer rises from 41% to 64% of the pool and distinct answers fall from 52 to 36; mean answer-choice surprisal drops from 1.80 to 1.58 bits. The tax is progressive: six of 44 models move individually (BH-FDR q=.10), all toward the mode, led by the most distinctive models, while the conformist floor is immobile. It is a sharpener, not a re-indexer -- the plain-chat modal answer survives in 28 of 31 categories. Defaults are register-indexed: a within-run re-sample (n=20) finds JSON shifts 53% of a model's stable chat defaults, mostly back to the crowd, and installs defaults absent from chat (Claude Fable 5 answers "cerulean" for colour 0% of the time in chat, 100% in JSON). Full-battery controls reveal a register gradient: compression is significant and specific to the answer-delivery formats models are trained to speak (JSON -0.22 bits, p=.0002; XML -0.19, p=.002), absent for YAML and CSV, and reversed for an arbitrary bracket wrapper (+0.13, p=.009) -- weighing the mechanism toward tool-use post-training. Enforcing the schema at the decoder (response_format) compresses no further than the request (-0.03 bits): the collapse lives in the model's response to the register, not the decoder. Structured output is how software consumes language models, and that surface is served by a measurably more homogeneous model than the chat surface on which models are evaluated, compared, and chosen.
| Comments: | 12 pages, 1 figure. Companion to the One-Word Census (arXiv:2607.12796). Code, data, and interactive explorer: this https URL |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.18476 [cs.CL] |
| (or arXiv:2607.18476v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.18476 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tapan Parikh [view email]
[v1]
Mon, 20 Jul 2026 19:47:16 UTC (85 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.