Do LLMs write like humans? Variation in grammatical and rhetorical styles
Abstract
Large language models (LLMs) are capable of writing grammatical text that follows instructions, answers questions, and solves problems. As they have advanced, it has become difficult to distinguish their output from human-written text. While past research has found some differences in features such as word choice and punctuation and developed classifiers to detect LLM output, none has studied the rhetorical styles of LLMs. Using several variants of Llama 3 and GPT-4o, we construct two parallel corpora of human- and LLM-written texts from common prompts. Using Douglas Biber’s set of lexical, grammatical, and rhetorical features, we identify systematic differences between LLMs and humans and between different LLMs. These differences persist when moving from smaller models to larger ones and are larger for instruction-tuned models than base models. This observation of differences demonstrates that despite their advanced abilities, LLMs struggle to match human stylistic variation. Attention to more advanced linguistic features can hence detect patterns in their behavior not previously recognized.
Article Details
Journal Info
Proceedings of the National Academy of Sciences
National Academy of Sciences
Authors (7)
Alex Reinhart
Department of Statistics and Data Science
Ben Markey
Department of English
Michael Laudenbach
Department of Humanities and Social Sciences
Kachatad Pantusen
Department of Statistics and Data Science
Ronald Yurko
Department of Statistics and Data Science
Gordon Weinberg
Department of Statistics and Data Science
David West Brown
Department of English