<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="https://emreview.ru/lib/pkp/xml/oai2.xsl" ?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/"
	xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
	xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/
		http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
	<responseDate>2026-08-11T03:36:21Z</responseDate>
	<request identifier="oai:emreview.ru:article/3130" metadataPrefix="jats" verb="GetRecord">https://emreview.ru/index.php/emr/oai</request>
	<GetRecord>
		<record>
			<header>
				<identifier>oai:emreview.ru:article/3130</identifier>
				<datestamp>2026-08-10T07:12:24Z</datestamp>
				<setSpec>emr:%D0%9F%D0%98</setSpec>
			</header>
			<metadata>
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="https://jats.nlm.nih.gov/publishing/1.1/" dtd-version="1.1" xsi:noNamespaceSchemaLocation="https://jats.nlm.nih.gov/archiving/1.4/xsd/JATS-archivearticle1.xsd" xml:lang="ru" specific-use="eps-0.1">
			<front>
			<journal-meta>
				<journal-id journal-id-type="publisher">emr</journal-id><journal-id journal-id-type="ojs">emr</journal-id>
				<journal-title-group>
			<journal-title xml:lang="ru">Управление образованием: теория и практика</journal-title><trans-title-group xml:lang="en"><trans-title>Education Management Review</trans-title></trans-title-group>
</journal-title-group>			<issn pub-type="epub">2311-2174</issn>			<publisher><publisher-name>Индивидуальный предприниматель Подколзин М.М.</publisher-name></publisher>
			<self-uri xlink:href="https://emreview.ru/index.php/emr"/>
		</journal-meta>
		<article-meta>
			<article-id pub-id-type="doi">10.25726/j2589-9037-6347-b</article-id><article-id pub-id-type="publisher-id">3130</article-id>
			<article-categories><subj-group subj-group-type="heading" xml:lang="en"><subject>APPLIED RESEARCH</subject></subj-group><subj-group subj-group-type="heading" xml:lang="ru"><subject>ПРИКЛАДНЫЕ ИССЛЕДОВАНИЯ</subject></subj-group></article-categories>
			<title-group><article-title xml:lang="ru">Стилистические девиации как маркер человеческого письма: сопоставительный анализ текстов на материале корпуса CoAT</article-title><trans-title-group xml:lang="en"><trans-title>Stylistic deviations as a marker of human writing: a comparative analysis of texts based on the CoAT corpus</trans-title></trans-title-group></title-group>
			<contrib-group content-type="author">
				<contrib contrib-type="author">
					<name-alternatives>
						<name name-style="western" specific-use="primary" xml:lang="ru">
							<surname>Медведева</surname>
							<given-names>Анастасия Вадимовна</given-names>
						</name>
						<name name-style="western" xml:lang="en">
							<surname>Medvedeva</surname>
							<given-names>Anastasia V.</given-names>
						</name>
					</name-alternatives>
					<xref ref-type="aff" rid="aff-1"/>
					<email>comma.j@mail.ru</email>
				</contrib>
			</contrib-group>
			<aff-alternatives id="aff-1">
				<aff xml:lang="ru"><institution content-type="orgname">Санкт-Петербургский политехнический университет Петра Великого, 195251, Санкт-Петербург, Политехническая улица, 29</institution></aff>
				<aff xml:lang="en"><institution content-type="orgname">Peter the Great St. Petersburg Polytechnic University, 195251, St. Petersburg, Politekhnicheskaya street, 29</institution></aff>
			</aff-alternatives>
			<pub-date date-type="collection"><year>2026</year></pub-date><pub-date date-type="pub" publication-format="epub">
				<day>30</day>
				<month>04</month>
				<year>2026</year>
			</pub-date>
			<volume seq="43">1616</volume>
			<issue>44</issue>
				<issue-id>130</issue-id><issue-title xml:lang="ru">Управление образованием: теория и практика</issue-title><issue-title xml:lang="en">Education Management Review</issue-title><fpage>498</fpage>
				<lpage>507</lpage>
			<history>
				<date date-type="received" iso-8601-date="2026-08-10">
					<day>10</day>
					<month>08</month>
					<year>2026</year>
				</date>
			</history>
			<permissions>
				<copyright-statement>Copyright (c) 2026 Управление образованием: теория и практика</copyright-statement>
				<copyright-year>2026</copyright-year>
				<copyright-holder>Управление образованием: теория и практика</copyright-holder>
				<license xml:lang="ru" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0">
					<license-p>Это произведение доступно по лицензии Creative Commons «Attribution-NonCommercial-NoDerivatives» («Атрибуция — Некоммерческое использование — Без производных произведений») 4.0 Всемирная.</license-p>
				</license>
				<license license-type="open-access" specific-use="metadata" xml:lang="ru" xlink:href="https://creativecommons.org/publicdomain/zero/1.0/">
					<license-p>Метаданные настоящей записи распространяются на условиях Creative Commons CC0 1.0 (передача в общественное достояние).</license-p>
				</license>
			</permissions>
			
			<self-uri xlink:href="https://emreview.ru/index.php/emr/article/view/3130"/>
			<abstract>В статье осуществляется сопоставительный лингвостилистический анализ русскоязычных текстов человеческого и машинного происхождения на материале корпуса CoAT, выступающего в качестве репрезентативного ресурса парных материалов, распределенных по новостному, энциклопедическому, сетевому, дневниковому и смешанному доменам. Центральное место отводится стилистическим девиациям, понимаемым как мотивированные отклонения от статистического стандарта, возникающие в пространстве между статистической и коммуникативной нормами и придающие высказыванию измерение индивидуально-авторской интенциональности, в противовес усредненному, вероятностно-оптимизированному письму, тяготеющему к нормативной гладкости и лишенному подлинного следа сознательного выбора. Анализ синтаксических параметров фиксирует у человеческих микротекстов большую среднюю длину предложения, составляющую 10,46 слова против 8,4 у сгенерированных, а также существенно более высокую частоту инверсий, достигающую 14% по сравнению с 4,7%, причем в неформальных жанрах, таких как дневники и социальные сети, машинные тексты демонстрируют полное отсутствие инверсивных конструкций, что указывает на принципиально различный механизм организации высказывания, обусловленный у человека коммуникативной задачей, а у модели – жесткостью жанровых шаблонов. Парцелляция, хотя и представлена близкими долями, в человеческих текстах выступает осмысленным интонационным приемом, тогда как в машинных чаще носит характер механического разрыва синтаксической конструкции. Пунктуационные характеристики выявляют сходство средних значений по тире и многоточиям при кардинально различном распределении и функциональной нагрузке: у человека эти знаки концентрируются в эмоционально-насыщенных контекстах, обеспечивая эффект паузации и смыслового акцента как проявление ответственного поступка, в то время как у машины они преимущественно воспроизводят формальные шаблоны, характерные для энциклопедических материалов, отражая отсутствие внутренней интенции и рефлексивной настройки на адресата. Особенно показательна крайне низкая воспроизводимость скобок в сгенерированных текстах, выступающих одним из наиболее тонких авторских маркеров. Лексические параметры демонстрируют наиболее выраженные расхождения по фактологической насыщенности: трехкратное преобладание чисел и дат в человеческих текстах свидетельствует о спонтанной связи письма с конкретным опытом и реальным миром, недоступной модели, оперирующей обобщенными вероятностными последовательностями и избегающей конкретики вне жанровой обязательности. Частота имен собственных обнаруживает меньшую дифференциацию, будучи в значительной степени детерминированной информационными жанрами, однако распределение по доменам дополнительно подчеркивает жанровую ограниченность машинного письма, не проявляющего фактографию вне обязательных контекстов. Полученные результаты подтверждают сущностный характер различий между текстом, рожденным сознанием, включенным в физический и социальный мир, и текстом, порожденным статистическим предсказанием следующего токена, что позволяет рассматривать стилистические девиации как перспективный маркер человеческого письма, обладающий высокой релевантностью для автоматизированных систем поиска и фильтрации информации, задач верификации контента, оценки качества генеративных моделей по критерию способности воспроизводить индивидуальную вариативность, противодействия дезинформации и развития цифровой лингвистики в целом. Качественный анализ распределения признаков по жанрам и категориям раскрывает, что количественные совпадения не отменяют фундаментального различия в природе отклонений – осознанного авторского выбора versus механической имитации, что открывает пути к созданию чувствительных типологий, адаптируемых для иных языков и жанров, и к лонгитюдному отслеживанию эволюции машинного идиостиля. Таким образом, работа предоставляет комплексное обоснование того, что формальный учет стилистических девиаций с учетом их качественной природы служит надежным инструментом разграничения человеческого и искусственного письма, способствуя как теоретическому осмыслению идиостиля в эпоху нейросетей, так и решению прикладных задач обработки естественного языка в условиях стремительного роста цифрового контента.</abstract><trans-abstract xml:lang="en">The article provides a comparative linguostylistic analysis of Russian-language texts of human and machine origin based on the CoAT corpus, which serves as a representative resource for paired materials distributed across news, encyclopedic, network, diary, and mixed domains. The central focus is on stylistic deviations, which are understood as motivated deviations from the statistical standard that occur in the space between the statistical and communicative norms and give the statement a dimension of individual authorial intentionality, as opposed to the average, probabilistically optimized writing that tends towards normative smoothness and lacks a genuine trace of conscious choice. The analysis of syntactic parameters reveals that human microtexts have a longer average sentence length of 10,46 words compared to 8,4 words in generated texts, as well as a significantly higher frequency of inversions, reaching 14% compared to 4,7%. In informal genres such as diaries and social media, machine-generated texts exhibit a complete absence of inversions, indicating a fundamentally different mechanism of sentence organization driven by human communication goals versus the rigidness of genre templates in the model. Parcellation, although represented by close fractions, is a meaningful intonation device in human texts, while in machine texts it is more often a mechanical break in the syntactic structure. Punctuation characteristics reveal similarities in the average values of dashes and dots, despite their drastically different distribution and functional load: In humans, these signs are concentrated in emotionally charged contexts, providing a pause effect and semantic emphasis as a manifestation of responsible action, while in machines, they predominantly reproduce formal templates characteristic of encyclopedic materials, reflecting the absence of internal intention and reflexive adjustment to the addressee. The extremely low reproducibility of brackets in generated texts, which are one of the most subtle authorial markers, is particularly noteworthy. Lexical parameters demonstrate the most pronounced discrepancies in terms of factual saturation: the three-fold predominance of numbers and dates in human texts indicates a spontaneous connection between writing and concrete experience and the real world, which is not available to a model that operates with generalized probabilistic sequences and avoids specificity outside of genre-specific requirements. The frequency of proper names shows less differentiation, as it is largely determined by information genres, but the distribution across domains further highlights the genre-specific limitations of machine writing, which does not exhibit factography outside of required contexts. The results obtained confirm the essential nature of the differences between text generated by a consciousness that is embedded in the physical and social world and text generated by the statistical prediction of the next token, which allows us to consider stylistic deviations as a promising marker of human writing that is highly relevant for automated information search and filtering systems, content verification tasks, evaluating the quality of generative models based on their ability to reproduce individual variability, countering disinformation, and the development of digital linguistics in general. A qualitative analysis of the distribution of features across genres and categories reveals that quantitative matches do not negate the fundamental difference in the nature of deviations – deliberate authorial choices versus mechanical imitation – which opens up avenues for creating sensitive typologies that can be adapted to other languages and genres, as well as for longitudinally tracking the evolution of machine idiostyle. Thus, the work provides a comprehensive justification for the fact that the formal account of stylistic deviations, taking into account their qualitative nature, serves as a reliable tool for distinguishing between human and artificial writing, contributing both to the theoretical understanding of idiostyle in the era of neural networks and to the solution of applied tasks of natural language processing in the context of the rapid growth of digital content.</trans-abstract><trans-abstract xml:lang="en"><p>The article provides a comparative linguostylistic analysis of Russian-language texts of human and machine origin based on the CoAT corpus, which serves as a representative resource for paired materials distributed across news, encyclopedic, network, diary, and mixed domains. The central focus is on stylistic deviations, which are understood as motivated deviations from the statistical standard that occur in the space between the statistical and communicative norms and give the statement a dimension of individual authorial intentionality, as opposed to the average, probabilistically optimized writing that tends towards normative smoothness and lacks a genuine trace of conscious choice. The analysis of syntactic parameters reveals that human microtexts have a longer average sentence length of 10,46 words compared to 8,4 words in generated texts, as well as a significantly higher frequency of inversions, reaching 14% compared to 4,7%. In informal genres such as diaries and social media, machine-generated texts exhibit a complete absence of inversions, indicating a fundamentally different mechanism of sentence organization driven by human communication goals versus the rigidness of genre templates in the model. Parcellation, although represented by close fractions, is a meaningful intonation device in human texts, while in machine texts it is more often a mechanical break in the syntactic structure. Punctuation characteristics reveal similarities in the average values of dashes and dots, despite their drastically different distribution and functional load: In humans, these signs are concentrated in emotionally charged contexts, providing a pause effect and semantic emphasis as a manifestation of responsible action, while in machines, they predominantly reproduce formal templates characteristic of encyclopedic materials, reflecting the absence of internal intention and reflexive adjustment to the addressee. The extremely low reproducibility of brackets in generated texts, which are one of the most subtle authorial markers, is particularly noteworthy. Lexical parameters demonstrate the most pronounced discrepancies in terms of factual saturation: the three-fold predominance of numbers and dates in human texts indicates a spontaneous connection between writing and concrete experience and the real world, which is not available to a model that operates with generalized probabilistic sequences and avoids specificity outside of genre-specific requirements. The frequency of proper names shows less differentiation, as it is largely determined by information genres, but the distribution across domains further highlights the genre-specific limitations of machine writing, which does not exhibit factography outside of required contexts. The results obtained confirm the essential nature of the differences between text generated by a consciousness that is embedded in the physical and social world and text generated by the statistical prediction of the next token, which allows us to consider stylistic deviations as a promising marker of human writing that is highly relevant for automated information search and filtering systems, content verification tasks, evaluating the quality of generative models based on their ability to reproduce individual variability, countering disinformation, and the development of digital linguistics in general. A qualitative analysis of the distribution of features across genres and categories reveals that quantitative matches do not negate the fundamental difference in the nature of deviations – deliberate authorial choices versus mechanical imitation – which opens up avenues for creating sensitive typologies that can be adapted to other languages and genres, as well as for longitudinally tracking the evolution of machine idiostyle. Thus, the work provides a comprehensive justification for the fact that the formal account of stylistic deviations, taking into account their qualitative nature, serves as a reliable tool for distinguishing between human and artificial writing, contributing both to the theoretical understanding of idiostyle in the era of neural networks and to the solution of applied tasks of natural language processing in the context of the rapid growth of digital content.</p></trans-abstract>
			
			
			<kwd-group xml:lang="en"><title>Keywords</title><kwd>stylistic deviations</kwd><kwd>generative neural networks</kwd><kwd>corpus linguistics</kwd><kwd>idiostyle</kwd><kwd>language norm</kwd><kwd>CoAT</kwd></kwd-group><kwd-group xml:lang="ru"><title>Ключевые слова</title><kwd>стилистические девиации</kwd><kwd>генеративные нейросети</kwd><kwd>корпусная лингвистика</kwd><kwd>идиостиль</kwd><kwd>языковая норма</kwd><kwd>CoAT</kwd></kwd-group><funding-group>
				<funding-statement xml:lang="ru">Исследование выполнено без внешнего финансирования.</funding-statement>
				<funding-statement xml:lang="en">The study was conducted without external funding.</funding-statement>
			</funding-group>
			<counts><page-count count="10"/></counts>
			<custom-meta-group><custom-meta><meta-name>issue-cover</meta-name><meta-value><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://emreview.ru/public/journals/1/cover_issue_130_ru_RU.png"/></meta-value></custom-meta></custom-meta-group><custom-meta-group>
				<custom-meta>
					<meta-name>metadata-license</meta-name>
					<meta-value><ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/publicdomain/zero/1.0/">CC0 1.0</ext-link></meta-value>
				</custom-meta>
			</custom-meta-group>
		</article-meta>
	</front>
	<back>
		<ref-list xml:lang="ru">
			<title>Список литературы</title>
			<ref id="R1"><mixed-citation>Абаева Е.С., Воеводина А.И. Идиостиль автора: вопросы параметризации // Филологические науки. Вопросы теории и практики. 2024. Т. 17. № 10. С. 3681-3687.</mixed-citation></ref>
			<ref id="R2"><mixed-citation>Бахтин М.М. Проблема речевых жанров // Эстетика словесного творчества. Москва: Искусство, 1979. С. 237-280.</mixed-citation></ref>
			<ref id="R3"><mixed-citation>Гальперин И.Р. Текст как объект лингвистического исследования. 3-е изд. Москва: УРСС, 2005. 137 с.</mixed-citation></ref>
			<ref id="R4"><mixed-citation>Горожанов А.И. Создание лингвистического корпуса на основе инструментов обработки естественного языка: планирование программных решений // Филологические науки. Вопросы теории и практики. 2023. Т. 16. № 5. С. 1616-1620.</mixed-citation></ref>
			<ref id="R5"><mixed-citation>Караулов Ю.Н. Русский язык и языковая личность. 7-е изд. Москва: ЛКИ, 2010. 264 с.</mixed-citation></ref>
			<ref id="R6"><mixed-citation>Клушина Н.И. Идиостиль в генеративном тексте // Коммуникативные исследования. 2025. Т. 12. № 4. С. 774-788.</mixed-citation></ref>
			<ref id="R7"><mixed-citation>Колмогорова А.В., Марголина А.В. Написанный vs сгенерированный текст: «естественность» как категория текстовая и психолингвистическая // Научный результат. Вопросы теоретической и прикладной лингвистики. 2024. Т. 10. № 2. С. 71-99.</mixed-citation></ref>
			<ref id="R8"><mixed-citation>Микаллеф Л.О. Лингвистика нейросетей как парадигма современной науки о языке // Мир науки, культуры, образования. 2025. № 1(110). С. 467-469.</mixed-citation></ref>
			<ref id="R9"><mixed-citation>Мордовин А.Ю. К вопросу о понятии репрезентативности корпуса текстов // Вестник ИГЛУ. 2009. № 1. С. 31-37.</mixed-citation></ref>
			<ref id="R10"><mixed-citation>Осетрова Е. В., Седова А. В. Характеристики сгенерированного текста: языковой и социально-коммуникативный анализ // Сибирский филологический форум. 2025. № 2 (31). С. 45–55. DOI: 10.24412/2587-7844-2025-2-45-55.</mixed-citation></ref>
			<ref id="R11"><mixed-citation>Прохоров А.И., Асадчая К.В. Инструментальные средства определения текста, сгенерированного при помощи нейросети // Научный вектор: сб. науч. тр. Под науч. ред. Е.Н. Макаренко. Т. 9. Ростов н/Д: Ростовский государственный экономический университет «РИНХ», 2023. С. 250-253.</mixed-citation></ref>
			<ref id="R12"><mixed-citation>Старкова Е.В. Проблема понимания феномена идиостиля в лингвистических исследованиях // Вестник Вятского государственного гуманитарного университета. 2015. № 5. С. 75-81.</mixed-citation></ref>
			<ref id="R13"><mixed-citation>Тельпов Р.Е., Ларцина С.В. Типовые различия естественных и сгенерированных нейронной сетью текстов в квантитативном аспекте // Научный диалог. 2023. Т. 12. № 7. С. 47-65.</mixed-citation></ref>
			<ref id="R14"><mixed-citation>Туркулец И.А. Композиционные особенности текстов, сгенерированных ChatGPT, как маркер несамостоятельности выполнения работ студентами // Правовая реальность в условиях цифровизации общества: материалы Всероссийской научно-практической конференции. Хабаровск: Дальневосточный государственный университет путей сообщения, 2023. С. 59-68.</mixed-citation></ref>
			<ref id="R15"><mixed-citation>Уразбаева Н.Ж. Человек и искусственный интеллект в письменной речи: проблема языковой интуиции нейросетей // Молодой ученый. 2026. № 16.1 (619.1). С. 24-25.</mixed-citation></ref>
			<ref id="R16"><mixed-citation>Черкасова М.Н., Тактарова А.В. Признаки сгенерированного текста в академическом дискурсе: проблема идентификации // Филологические науки. Вопросы теории и практики. 2024. Т. 17. № 7. С. 2226-2232.</mixed-citation></ref>
			<ref id="R17"><mixed-citation>Shamardina T., Saidov M., Fenogenova A., Tumanov A., Zemlyakova A., Lebedeva A., Gryaznova E., Shavrina T., Mikhailov V., Artemova E. CoAT: Corpus of artificial texts // Natural Language Processing. 2025. Vol. 31. Issue 1. P. 150–175. DOI: 10.1017/nlp.2024.38.</mixed-citation></ref>
		</ref-list>
	</back>
</article>			</metadata>
		</record>
	</GetRecord>
</OAI-PMH>
