Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning
Quick Answer
The paper introduces Semi-CoT, a semi-supervised learning framework leveraging unlabeled questions to generate pseudo reasoning chains for large language models.
Quick Take
Experiments on benchmarks like AQuA and GSM8K show pseudo-answer precision between 91.36% and 100%, indicating potential for effective reasoning signal generation, though challenges remain in demonstration selection.
Key Points
- Semi-CoT constructs pseudo reasoning supervision from unlabeled questions.
- Pilot experiments show pseudo-answer precision ranging from 91.36% to 100%.
- Entropy gate effectively selects high-precision pseudo-CoTs.
- Results indicate potential but highlight challenges in demonstration selection.
- Negative transfer observed on AQuA; MultiArith performance reached a ceiling.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Chain-of-thought (CoT) reasoning has emerged as an effective approach for activating latent reasoning capabilities in large language models. However, most existing CoT methods use reasoning chains mainly as inference-time prompts, while the generated reasoning traces are rarely reused as semi-supervised learning signals. In this report, we define \textbf{Semi-supervised Chain-of-Thought Learning} and propose \textbf{Semi-CoT}, a simple framework that uses unlabeled questions to construct pseudo reasoning supervision. Semi-CoT samples multiple pseudo-CoTs for each unlabeled question, estimates answer-level semantic entropy, and selects low-entropy reasoning chains as reliable pseudo-CoT demonstrations. This extends the self-training view of CoT from inference-time refinement to semi-supervised pseudo-supervision. Pilot experiments on AQuA, SVAMP, GSM8K, and MultiArith show that the entropy gate selects high-precision pseudo-CoTs, with pseudo-answer precision ranging from $91.36\%$ to $100\%$. Semi-CoT also gives small gains on SVAMP and GSM8K, while AQuA shows negative transfer and MultiArith reaches a ceiling. These results suggest that unlabeled questions can provide reliable pseudo reasoning signals, but their effective use still requires stronger demonstration selection or student training.
| Comments: | Tech Report |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.01511 [cs.AI] |
| (or arXiv:2607.01511v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.01511 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hongyang He [view email]
[v1]
Wed, 1 Jul 2026 22:17:39 UTC (18,967 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.