APPLIED RESEARCH

Stylistic deviations as a marker of human writing: a comparative analysis of texts based on the CoAT corpus

Authors

  • Anastasia V. Medvedeva Peter the Great St. Petersburg Polytechnic University, 195251, St. Petersburg, Politekhnicheskaya street, 29

How to cite

GOST Medvedeva A. V. Stylistic deviations as a marker of human writing: a comparative analysis of texts based on the CoAT corpus // Education Management Review. 2026. Vol. 16. No. 4. P. 498-507. DOI: 10.25726/j2589-9037-6347-b
APA Medvedeva, A. V. (2026). Stylistic deviations as a marker of human writing: a comparative analysis of texts based on the CoAT corpus. Education Management Review, 16(4), 498-507. https://doi.org/10.25726/j2589-9037-6347-b

Abstract

The article provides a comparative linguostylistic analysis of Russian-language texts of human and machine origin based on the CoAT corpus, which serves as a representative resource for paired materials distributed across news, encyclopedic, network, diary, and mixed domains. The central focus is on stylistic deviations, which are understood as motivated deviations from the statistical standard that occur in the space between the statistical and communicative norms and give the statement a dimension of individual authorial intentionality, as opposed to the average, probabilistically optimized writing that tends towards normative smoothness and lacks a genuine trace of conscious choice. The analysis of syntactic parameters reveals that human microtexts have a longer average sentence length of 10,46 words compared to 8,4 words in generated texts, as well as a significantly higher frequency of inversions, reaching 14% compared to 4,7%. In informal genres such as diaries and social media, machine-generated texts exhibit a complete absence of inversions, indicating a fundamentally different mechanism of sentence organization driven by human communication goals versus the rigidness of genre templates in the model. Parcellation, although represented by close fractions, is a meaningful intonation device in human texts, while in machine texts it is more often a mechanical break in the syntactic structure. Punctuation characteristics reveal similarities in the average values of dashes and dots, despite their drastically different distribution and functional load: In humans, these signs are concentrated in emotionally charged contexts, providing a pause effect and semantic emphasis as a manifestation of responsible action, while in machines, they predominantly reproduce formal templates characteristic of encyclopedic materials, reflecting the absence of internal intention and reflexive adjustment to the addressee. The extremely low reproducibility of brackets in generated texts, which are one of the most subtle authorial markers, is particularly noteworthy. Lexical parameters demonstrate the most pronounced discrepancies in terms of factual saturation: the three-fold predominance of numbers and dates in human texts indicates a spontaneous connection between writing and concrete experience and the real world, which is not available to a model that operates with generalized probabilistic sequences and avoids specificity outside of genre-specific requirements. The frequency of proper names shows less differentiation, as it is largely determined by information genres, but the distribution across domains further highlights the genre-specific limitations of machine writing, which does not exhibit factography outside of required contexts. The results obtained confirm the essential nature of the differences between text generated by a consciousness that is embedded in the physical and social world and text generated by the statistical prediction of the next token, which allows us to consider stylistic deviations as a promising marker of human writing that is highly relevant for automated information search and filtering systems, content verification tasks, evaluating the quality of generative models based on their ability to reproduce individual variability, countering disinformation, and the development of digital linguistics in general. A qualitative analysis of the distribution of features across genres and categories reveals that quantitative matches do not negate the fundamental difference in the nature of deviations – deliberate authorial choices versus mechanical imitation – which opens up avenues for creating sensitive typologies that can be adapted to other languages and genres, as well as for longitudinally tracking the evolution of machine idiostyle. Thus, the work provides a comprehensive justification for the fact that the formal account of stylistic deviations, taking into account their qualitative nature, serves as a reliable tool for distinguishing between human and artificial writing, contributing both to the theoretical understanding of idiostyle in the era of neural networks and to the solution of applied tasks of natural language processing in the context of the rapid growth of digital content.

Keywords

stylistic deviations generative neural networks corpus linguistics idiostyle language norm CoAT

References

Абаева Е.С., Воеводина А.И. Идиостиль автора: вопросы параметризации // Филологические науки. Вопросы теории и практики. 2024. Т. 17. № 10. С. 3681-3687.

Бахтин М.М. Проблема речевых жанров // Эстетика словесного творчества. Москва: Искусство, 1979. С. 237-280.

Гальперин И.Р. Текст как объект лингвистического исследования. 3-е изд. Москва: УРСС, 2005. 137 с.

Горожанов А.И. Создание лингвистического корпуса на основе инструментов обработки естественного языка: планирование программных решений // Филологические науки. Вопросы теории и практики. 2023. Т. 16. № 5. С. 1616-1620.

Караулов Ю.Н. Русский язык и языковая личность. 7-е изд. Москва: ЛКИ, 2010. 264 с.

Клушина Н.И. Идиостиль в генеративном тексте // Коммуникативные исследования. 2025. Т. 12. № 4. С. 774-788.

Колмогорова А.В., Марголина А.В. Написанный vs сгенерированный текст: «естественность» как категория текстовая и психолингвистическая // Научный результат. Вопросы теоретической и прикладной лингвистики. 2024. Т. 10. № 2. С. 71-99.

Микаллеф Л.О. Лингвистика нейросетей как парадигма современной науки о языке // Мир науки, культуры, образования. 2025. № 1(110). С. 467-469.

Мордовин А.Ю. К вопросу о понятии репрезентативности корпуса текстов // Вестник ИГЛУ. 2009. № 1. С. 31-37.

Осетрова Е. В., Седова А. В. Характеристики сгенерированного текста: языковой и социально-коммуникативный анализ // Сибирский филологический форум. 2025. № 2 (31). С. 45–55. DOI: 10.24412/2587-7844-2025-2-45-55.

Прохоров А.И., Асадчая К.В. Инструментальные средства определения текста, сгенерированного при помощи нейросети // Научный вектор: сб. науч. тр. Под науч. ред. Е.Н. Макаренко. Т. 9. Ростов н/Д: Ростовский государственный экономический университет «РИНХ», 2023. С. 250-253.

Старкова Е.В. Проблема понимания феномена идиостиля в лингвистических исследованиях // Вестник Вятского государственного гуманитарного университета. 2015. № 5. С. 75-81.

Тельпов Р.Е., Ларцина С.В. Типовые различия естественных и сгенерированных нейронной сетью текстов в квантитативном аспекте // Научный диалог. 2023. Т. 12. № 7. С. 47-65.

Туркулец И.А. Композиционные особенности текстов, сгенерированных ChatGPT, как маркер несамостоятельности выполнения работ студентами // Правовая реальность в условиях цифровизации общества: материалы Всероссийской научно-практической конференции. Хабаровск: Дальневосточный государственный университет путей сообщения, 2023. С. 59-68.

Уразбаева Н.Ж. Человек и искусственный интеллект в письменной речи: проблема языковой интуиции нейросетей // Молодой ученый. 2026. № 16.1 (619.1). С. 24-25.

Черкасова М.Н., Тактарова А.В. Признаки сгенерированного текста в академическом дискурсе: проблема идентификации // Филологические науки. Вопросы теории и практики. 2024. Т. 17. № 7. С. 2226-2232.

Shamardina T., Saidov M., Fenogenova A., Tumanov A., Zemlyakova A., Lebedeva A., Gryaznova E., Shavrina T., Mikhailov V., Artemova E. CoAT: Corpus of artificial texts // Natural Language Processing. 2025. Vol. 31. Issue 1. P. 150–175. DOI: 10.1017/nlp.2024.38.

Issue

Section

APPLIED RESEARCH

Metrics

0 views
0 downloads
Want to publish with us?
Submit an article

Machine-readable metadata