HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
Quick Answer
HSS-Synth introduces a novel data synthesis pipeline for humanities and social sciences, generating 237k high-quality instruction-tuning samples that outperform 14 leading baselines across 16 benchmarks.
Quick Take
The fine-tuned Qwen3-8B-Base sets a new state-of-the-art in performance, enhancing human preference and knowledge capabilities without performance fluctuations.
Key Points
- HSS-Synth covers 14 mainstream fields in humanities and social sciences.
- The pipeline uses multi-step filtering and text refinement for seed document construction.
- It employs teacher-forced answering to anchor semantics and reduce hallucinations.
- The generated samples outperform existing models on 16 benchmarks.
- Code for HSS-Synth is publicly available for further research.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Ru Peng, Tianyu Zhao, Xijun Gu, Zhiting Fan, Haokai Xu, Jinyang Zhang, Yawen Zeng, Yihong Zhuang, Kexin Yang, Junyang Lin, Dayiheng Liu, Junbo Zhao
Abstract:High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly. Data synthesis is a viable alternative and succeeds on closed tasks, yet the humanities and social sciences (HSS) are overlooked, and their open-ended nature makes synthesis challenging. Moving beyond prior capability-centric, fragmented attempts, we adopt a subject-centric paradigm, define the first HSS domain system covering 14 mainstream fields, and introduce HSS-Synth, the first data synthesis pipeline for HSS. HSS-Synth comprises: (1) constructing seed documents from web corpora via multi-step filtering and text refinement evaluated by a judge; (2) specifying "requirements + persona" to backtranslate seed documents into diverse yet faithful instructions with a strict Q&A alignment check; and (3) breaking LLM response limits via teacher-forced Answering that feeds seed documents during response generation to anchor semantics, reduce hallucinations, and preserve tone and integrity. HSS-Synth yields 237k high-quality, diverse instruction-tuning samples that outperform 14 leading baselines on 16 benchmarks. The fine-tuned Qwen3-8B-Base sets a new SOTA and approaches the official Qwen3-8B, improving both human preference and knowledge capabilities without performance seesaws. Extensive experiments demonstrate HSS-Synth's robustness and transferability. Our code is publicly available at this https URL.
| Comments: | ACL Findings 2026 Paper |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.27379 [cs.CL] |
| (or arXiv:2607.27379v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.27379 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ru Peng [view email]
[v1]
Wed, 29 Jul 2026 18:37:14 UTC (16,109 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.