Distinguishing Artificial from Authentic: Evaluating LLMs for Detecting LLM-Generated Content
Quick Answer
This study evaluates the ability of large language models (LLMs) to detect their own generated content across various educational tasks.
Quick Take
Findings reveal that detection accuracy varies significantly by task type, with better performance in programming exercises compared to short-answer questions. The research underscores the limitations of relying solely on for identifying AI-generated student work.
Key Points
- LLMs show high detection accuracy for programming tasks but struggle with short-answer questions.
- Prompt design and response length significantly affect detection performance in reflective writing.
- LLMs often misjudge their own outputs as more human-like than authentic student responses.
- Detection effectiveness varies greatly across different educational task types.
- Caution is advised when using LLMs as standalone tools for identifying AI-generated work.
DeepSignal Analysis
What happened
A study evaluated the ability of large language models (LLMs) to detect their own generated content across various educational tasks. The results indicated that detection accuracy is highly dependent on the task type, with programming exercises yielding better results than short-answer questions.
Key evidence
- The study found that LLMs performed better in detecting their outputs in programming tasks compared to reflective writing and short-answer questions.
- Detection accuracy varied significantly based on factors such as prompt design and response length, particularly affecting reflective writing tasks.
- LLMs often misjudged their own outputs as more human-like than actual student responses, especially in short-answer questions.
Why it matters
The findings highlight the limitations of using LLMs as standalone tools for identifying AI-generated student work. While LLMs can be effective in certain contexts, their reliability varies significantly across different educational tasks, raising concerns about their use in academic integrity assessments.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:As large language models (LLMs) are increasingly used by students to generate natural language responses and program code, there is growing interest in whether LLMs themselves can be used to distinguish AI-generated work from human-authored submissions. In this paper, we investigate the extent to which LLMs can detect their own generated content across multiple educational task types, including programming exercises, reflective writing, and short-answer questions. Using authentic student responses and multiple variants of LLM-generated answers, we evaluate detection performance under different prompting strategies and output formats. Our study addresses three research questions: (1) how accurately LLMs can identify their own outputs across task domains, (2) how detection effectiveness is influenced by factors such as prompt design, response length, and task type, and (3) what characteristics of LLM-generated responses contribute to successful or failed detection. Our findings show that LLM-based detection is highly task-dependent: detection is substantially more reliable for programming tasks and longer reflective responses, but performs poorly for short-answer questions, where LLMs frequently judge their own outputs as more human-like than authentic student responses. We further find that prompt framing and response verbosity have a pronounced effect on detectability in reflective writing tasks, with relatively minor prompt variations significantly reducing detection accuracy, while programming-related detection is more robust to prompt changes. Together, these results highlight both the potential and the limitations of LLM self-detection in educational settings and suggest caution in relying on LLMs as standalone tools for identifying AI-generated student work.
| Comments: | 8 pages, 5 figures, 3 tables |
| Subjects: | Computation and Language (cs.CL); Computers and Society (cs.CY) |
| Cite as: | arXiv:2607.20446 [cs.CL] |
| (or arXiv:2607.20446v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.20446 arXiv-issued DOI via DataCite |
Submission history
From: Juho Leinonen [view email]
[v1]
Wed, 13 May 2026 13:38:40 UTC (1,029 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.