Parametric optimization of input data array construction based on character cross-links
DOI:
https://doi.org/10.34121/1028-9763-2026-3-72-81Keywords:
attribution of authorship, program code, code refactoring, input data array, symbolic crosslinks, informative sufficiency limit, microservices programming, parametric optimization, data monitoring and analysis, machine learning, web application development in Java, generative models, web programming, server programmingAbstract
The problem of parametric optimization of the input data array (IDA) construction process for source code authorship attribution is considered, in particular for distinguishing between human-written and large-language-model-generated code. A language-independent feature type has been further developed, namely character-level cross-links between words, represented as ordered pairs of word prefixes and suffixes within a sliding window. By considering all pairs of words, rather than only adjacent ones, these features are insensitive to local rearrangements of code fragments. The constructed input data array is evaluated using a tournament of 39 machine learning model architectures, and the classification quality is assessed both at the level of individual windows and at the author level by majority voting. A full-factorial computational experiment of 168 trials has been carried out for 12 authors (10 humans and 2 generative models) in four programming languages. The effects of window size, the informational sufficiency threshold, the filtering mode, and the cross-link rank on the proportion of correctly classified objects and the size of the feature dictionary have been studied. It has been established that cross-links improve performance in 37 % of the evaluated configurations. The greatest improvement (up to +7.56 %) is observed for a window size of 2,000 characters, which represents the point of maximum contribution of the additional features rather than the point of highest overall performance. The Pareto-optimal configuration uses a 500-character window, a threshold of three and the averaging filtering mode, achieving 99,37 % of correctly classified objects with a dictionary of only 310 features, whereas the maximum mode increases quality at the cost of a nearly twentyfold growth of the dictionary. The limits of applicability of the method and the conditions under which cross-links reduce classification quality are determined, in particular for authors with a high baseline due to the ceiling effect. Таbl.: 5. Figs.: 4. Refs.: 9 titles.
References
1. Голуб М.С. Формування масиву чисельних ознак для класифікації україномовних текстів в інформаційній технології інтелектуального моніторингу: дис. канд. техн. наук: 05.13.06. Черкаси: ЧДТУ, 2018. 157 с.
2. Caliskan-Islam A., Harang R., Liu A. et al. De-anonymizing programmers via code stylometry. Proceedings of the 24th USENIX Security Symposium. Washington, 2015. P. 255–270.
3. Abuhamad M., AbuHmed T., Mohaisen A., Nyang D. Large-scale and language-oblivious code authorship identification. Proc. of the 2018 ACM SIGSAC. Conference on Computer and Communications Security. Toronto, 2018. P. 101–114.
4. Li Z., Chen G.Q., Chen C. et al. Towards robustness of deep program processing models — detection, estimation and enhancement (RoPGen). Proc. of the 44th International Conference on Software Engineering (ICSE). Pittsburgh, 2022. P. 1097–1108.
5. Pan W.H., Chok M.J., Wong J.L.S. et al. Assessing AI detectors in identifying AI-generated code. Proc. of the 46th International Conference on Software Engineering: Software Engineering in Society (ICSESEET). Lisbon, 2024. P. 1–12.
6. Nguyen P.T., Di Rocco J., Di Ruscio D. et al. GPTSniffer: A CodeBERT-based classifier to detect source code written by ChatGPT. Journal of Systems and Software. 2024. Vol. 214. P. 112059.
7. Oedingen M., Engelhardt R.C., Denz R. et al. ChatGPT code detection: Techniques for uncovering the source of code. AI. 2024. Vol. 5, N 3. P. 1066–1094.
8. Ивахненко А.Г. Индуктивный метод самоорганизации моделей сложных систем. Киев: Наукова думка, 1981. 296 с.
9. Голуб С.В., Немов Р.Г. Формування масиву чисельних ознак для класифікації авторства програмних кодів із використанням символьних крос-зв’язків між словами. Математичні машини і системи. 2026. № 2. С. 70–78. DOI: https://doi.org/10.34121/1028-9763-2026-2-70-78.

