Formation of an array of numerical features for classifying the authorship of software codes using symbolic cross-links between words
DOI:
https://doi.org/10.34121/1028-9763-2026-2-70-78Keywords:
code attribution, text classification, feature dictionary, input data array, information sufficiency threshold, cross-links between words, intellectual monitoring, GMDHAbstract
The paper addresses the problem of automated source-code authorship attribution as a component of the information technology for intellectual monitoring. Existing approaches to code attribution rely mainly on language-specific syntactic features — abstract syntax trees, tokens, or lexical constructs — and consequently do not generalize across programming languages, whereas real-world developers frequently write code in multiple languages while retaining a recognizable individual style. As an alternative, the methodology for constructing an input data array (IDA) within the S.V. Holub research school, developed in the dissertation research of M.S. Holub for the classification of Ukrainian-language texts, which employs a probabilistic feature-informativeness criterion and an information sufficiency threshold (IST), has been used. A new feature type — symbolic cross-links between words — that extends the Holub feature dictionary by capturing paired combinations of prefixes and suffixes of code identifiers within a fixed-length window, is used. Cross-links of three ranks (1×1, 2×2, 3×3) are formalized as occurrence frequencies of ordered k-character string pairs over all ordered word pairs within a window. The effectiveness of the proposed approach has been experimentally investigated on a dataset of 12 authors (10 human developers and 2 generative artificial-intelligence models — ChatGPT and Claude Code) across four programming languages (Java, JavaScript, TypeScript, and Python), comprising 119 source classes and 641 windows. In the within-language scenario, 89.8–100 % of windows have been correctly classified. In the cross-language scenario, the proposed method achieves 100 % accuracy in window classification at a window size of 500 characters, corresponding to a method advantage of up to +1.30 %. Cross-link features actively pass the IST and constitute up to 77 % of the adaptive dictionary volume, which demonstrates their high informativeness as a new, language-independent type of features for problems of software code authorship attribution. Таbl.: 3. Refs.: 12 titles.
References
1. Голуб М.С. Формування масиву вхідних даних при класифікації текстів у технології інформаційного моніторингу. Математичні машини і системи. 2018. № 1. С. 59–66.
2. Голуб М.С. Формування масиву чисельних ознак для класифікації україномовних текстів в інформаційній технології інтелектуального моніторингу: дис. канд. техн. наук: 05.13.06 / Черкаський державний технологічний університет. Черкаси, 2018. 137 с.
3. Ивахненко А.Г. Индуктивный метод самоорганизации моделей сложных систем. Киев: Наукова думка, 1981. 296 с.
4. Голуб С.В., Жирякова І.А., Куницька С.Ю., Авраменко В.П. Методи розвитку моніторингових інтелектуальних систем. Інформація, комунікація, суспільство 2019: матеріали 8-ї Міжнар. наук. конф. ICS-2019. Львів: Видавництво Львівської політехніки, 2019. С. 65–67.
5. Голуб М.С. Дисперсійний метод формування точок спостереження в інформаційній технології класифікації текстів. Вісник інженерної академії України. 2017. № 3. С. 38–42.
6. Немов Р.Г., Голуб С.В. Агентне програмування інтелектуального аналізу кодів програм. 13 міжнародна наукова конференція ІКС-2024. Львів, 2024. С. 133–135.
7. Немов Р.Г., Голуб С.В., Немченко В.В. Структурна динаміка програмного агента інформаційного моніторингу. ІТСМ. Івано-Франківськ, 2023. С. 91–94.
8. Голуб М.С. Вибір ознак у процесі інтелектуальної обробки текстових повідомлень. Інформація, комунікація, суспільство 2014: матеріали 3-ї Міжнар. наук. конф. ICS-2014. Львів: Видавництво Львівської політехніки, 2014. С. 148–149.
9. Голуб С.В., Константиновська О.В., Голуб М.С. Формування показників масиву вхідних даних для ідентифікації авторства текстових повідомлень. Системи обробки інформації: зб. наук. праць. Харків: Харківський університет повітряних сил імені Івана Кожедуба, 2014. Вип. 2 (118). С. 89–92.
10. Голуб М.С. Формування словника ознак для класифікації україномовних текстів в інформаційній технології багаторівневого інтелектуального моніторингу. Інформація, комунікація, суспільство 2019: матеріали 8-ї Міжнар. наук. конф. ICS-2019. Львів: Видавництво Львівської політехніки, 2019. С. 68–70.
11. Голуб С.В., Мартинова Г.І., Голуб М.С. Моделювання діалектного тексту в технології багаторівневого інформаційного моніторингу. Математичні машини і системи. 2016. № 4. С. 76-83.
12. Немов Р.Г., Голуб С.В. Агентне програмування інтелектуального аналізу кодів програм. I Міжнародна науково-практична конференція. Харків-Яремче, 2025. С. 218–220.

