A Systematic Analysis of Linguistic Features in AI-Generated Text Detection Across Domains and Models
Quick Answer
This paper shows that A large-scale study reveals that 284 linguistic features can effectively distinguish AI-generated text from human-written text across 27 LLMs and ten domains.
Quick Take
While many indicators are context-dependent, measures of lexical richness consistently serve as robust signals, enhancing interpretability for non-experts.
Key Points
- Study assesses 284 linguistic features across 27 and ten text domains.
- Classifiers based on linguistic features reliably distinguish AI and human text.
- Lexical richness measures are robust across different models and domains.
- Findings address gaps in understanding AI-generated text characteristics.
- Results support more reliable analyses of AI-generated language.
Paper Resources
Article Excerpt
From source RSS / original summaryarXiv:2606. 04177v1 Announce Type: new Abstract: Interpretable linguistic features offer a promising approach for explaining why a given text appears machine-generated, particularly for non-expert users. However, existing findings on which features reliably indicate -generated text remain fragmented across feature sets, models, and text domains. To address this gap, we conduct a large-scale empirical study assessing the robustness of linguistic signals for characterizing AI-generated text.
Our analysis covers 284 interpretable linguistic features across outputs from 27 LLMs and ten text domains under cross-model and cross-domain generalization settings. We show that classifiers based solely on linguistic features can reliably distinguish AI-generated from human-written text. However, many previously proposed indicators prove strongly context-dependent, with the exception of measures of lexical richness, which remain robust signals across model families and text domains.
These results demonstrate which linguistic signals generalize across contexts and provide a foundation for more reliable, interpretable analyses of AI-generated language.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.