Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models
Quick Answer
This paper shows that The Narration-of-Thought (NoT) system prompt significantly enhances ethical reasoning in large language models, reducing stakeholder collapse from 31% to under 1% and uncertainty suppression from 72% to 1-24% across four model generators.
Quick Take
This method requires no additional training and achieves a consensus increase from 6% to 95% in multi-stakeholder debates, providing a robust framework for ethical decision-making.
Key Points
- NoT organizes ethical reasoning into five sections: protagonist, stakeholders, consequences, uncertainty, commitment.
- Achieved a stakeholder collapse reduction from 31% to under 1% across 100 DailyDilemmas scenarios.
- Uncertainty suppression decreased from up to 72% to 1-24% across all models tested.
- Extended to a five-round debate, achieving 95% consensus from a 6% standoff.
- NoT requires no additional training, parameters, or fine-tuning.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with a stake in the outcome) and uncertainty suppression (no explicit unknowns or hedges before committing to an action). We introduce narration-of-thought (NoT), a system prompt that structures chain-of-thought into five sections: protagonist, stakeholders, two-step consequences, uncertainty, then commitment. NoT adds no training, parameters, or fine-tuning. On 100 DailyDilemmas scenarios across four generators from three vendors, NoT cuts stakeholder collapse from up to 31% to under 1% and uncertainty suppression from up to 72% to 1-24% on every model. A matched-budget verbose-CoT control rules out token spend as the active ingredient; NoT retains Cliff's delta advantages of +0.79 to +0.90 on stakeholder count and +0.65 to +0.93 on uncertainty score for three of four generators, and a section ablation attributes each shift to its specific sub-instruction. Textual-gradient descent initialised at NoT improves the scaffold further; a cross-family training judge (different vendor from the generator) dominates an in-family one on every measured axis. Extended to a five-round multi-stakeholder debate protocol, the scaffold converts a 6% standoff into 95% full consensus on a calibration set and 100% combined convergence on a DailyDilemmas replication. The resulting traces externalise the stakeholders, consequences, and uncertainty grounding each commitment, providing an auditable substrate for dependable agentic deployment.
| Comments: | 24 pages, 8 figures, 16 tables. To appear at ACL 2026 (submitted via ARR) |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY) |
| ACM classes: | I.2.7; I.2.1; K.4.1 |
| Cite as: | arXiv:2606.26366 [cs.AI] |
| (or arXiv:2606.26366v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.26366 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Patrick Cooper [view email]
[v1]
Wed, 24 Jun 2026 20:23:16 UTC (340 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.