Preprint
Large Language Models

A large-scale comparison of human-written versus ChatGPT-generated essays

Steffen Herbold(University of Passau), Annette Hautli-Janisz(University of Passau), Ute Heuer(University of Passau), Zlata Kikteva(University of Passau), Alexander Trautsch(University of Passau)
October 30, 2023Scientific Reports344 citations

344

Citations

19

Influential Citations

Scientific Reports

Venue

2023

Year

Abstract

ChatGPT and similar generative AI models have attracted hundreds of millions of users and have become part of the public discourse. Many believe that such models will disrupt society and lead to significant changes in the education system and information generation. So far, this belief is based on either colloquial evidence or benchmarks from the owners of the models-both lack scientific rigor. We systematically assess the quality of AI-generated content through a large-scale study comparing human-written versus ChatGPT-generated argumentative student essays. We use essays that were rated by a large number of human experts (teachers). We augment the analysis by considering a set of linguistic characteristics of the generated essays. Our results demonstrate that ChatGPT generates essays that are rated higher regarding quality than human-written essays. The writing style of the AI models exhibits linguistic characteristics that are different from those of the human-written essays. Since the technology is readily available, we believe that educators must act immediately. We must re-invent homework and develop teaching concepts that utilize these AI models in the same way as math utilizes the calculator: teach the general concepts first and then use AI tools to free up time for other learning objectives.

Analysis

Why This Paper Matters

This paper addresses a critical gap in the ongoing debate about generative AI's impact on education. While anecdotal claims and proprietary benchmarks abound, rigorous scientific evidence on the quality of AI-generated content has been scarce. By systematically comparing human-written and ChatGPT-generated essays using teacher evaluations, the study provides concrete, actionable insights for educators, policymakers, and AI developers. The finding that ChatGPT essays are rated higher in quality than human-written ones challenges assumptions about AI's limitations and underscores the urgency for educational reform.

The research is particularly timely as generative AI tools become widely accessible. The authors' call to "re-invent homework" and adopt AI in teaching—analogous to calculators in math—offers a constructive path forward. This paper moves beyond fear-mongering or hype, grounding recommendations in empirical data.

Technical Contributions

  • Large-scale empirical comparison: The study uses a substantial dataset of essays rated by many human experts, providing statistical power and ecological validity.
  • Linguistic feature analysis: Beyond quality ratings, the authors examine specific linguistic characteristics (e.g., vocabulary richness, syntactic complexity) that differentiate AI from human writing.
  • Teacher-based evaluation: Using actual educators as raters ensures relevance to real-world educational contexts, unlike automated metrics.
  • Reproducible methodology: The approach can be extended to other AI models, languages, or essay types.

Results

  • ChatGPT-generated essays received higher quality ratings from teachers compared to human-written essays.
  • The writing style of ChatGPT exhibits distinct linguistic characteristics (e.g., different patterns in word choice, sentence structure) relative to human essays.
  • These differences suggest that AI-generated text can be identified through stylistic analysis, though the quality advantage persists.

Significance

This paper has immediate implications for education: it provides evidence that AI can produce high-quality written content, necessitating a shift in how student work is assessed and assigned. For the AI field, it demonstrates that current LLMs can match or exceed human performance in structured writing tasks, raising the bar for evaluation benchmarks. The work also highlights the need for new teaching paradigms that leverage AI as a tool rather than viewing it as a threat. Future research can build on this by exploring other genres, multilingual settings, and longitudinal impacts on learning outcomes.