Machine learning of multilayer models of a monitoring software agent
DOI:
https://doi.org/10.34121/1028-9763-2025-2-76-95Keywords:
intelligent monitoring, software agent, multilayer modelling, machine learningAbstract
The use of an agent-based approach to the implementation of information technology for intelligent monitoring allows for obtaining special forms of multi-agent systems (MAS) to support decision-making in various subject areas. The structure of MAS for intelligent monitoring is formed by a set of software agents for various purposes. This paper presents the results of research on one of the MAS elements ― a monitoring software agent. In this system, it ensures the implementation of typical intellectual tasks of classification, identification, forecasting, and others. In the context of crisis monitoring, to obtain a useful agent model, methods of increasing the diversity of an agent-based model synthesizer (AMS), in particular, multilayer modelling, are used. Models with a heterogeneous structure and the same modelled indicator are combined into layers. Recirculation is one of the popular methods of multilayer modelling. The signal from the model output is added to the input data set (IDS) as an additional feature, and the IDS is sent to the input of the same AMS. The agent-based model synthesizer determines which AMS will be used adaptively to the properties of each subsequent input data set. Before synthesizing the model, each of the existing AMS is tested, and the best one is selected according to the specified criteria. If it is not possible to build a useful model, the synthesizer will adjust the AMS with a more complex structure. Improving the processes of adaptive construction of multilayer model ensembles using the recirculation method allows us to expand the capabilities of agent-based model synthesizers. This paper presents the results of applying machine learning methods to build algorithms for the synthesis of multilayer models using the recirculation method and strategies for the adaptive selection of the best AMS. The increase in the accuracy of modelling results has been experimentally confirmed. The obtained results allow us to build rules for the behaviour of a software agent and formulate requirements for MAS software.
References
1. Голуб С.В., Толбатов Д.В. Удосконалення методу синтезу багатошарових моделей моніторингового програмного агента. Актуальні проблеми автоматизації та інформаційних технологій. Дніпро, 2023. Т. 27. С. 53–60. DOI: https://doi.org/10.15421/432306.94 ISSN 1028-9763. Математичні машини і системи. 2025. № 2
2. Borisov V., Leemann T., Seßler K., Haug J., Pawelczyk M., Kasneci G. Deep Neural Networks and Tabular Data: A Survey. IEEE Trans. Neural Netw. Learn. Syst. 2022. Vol. 35, N 6. Р. 7499–7519. DOI: https://doi.org/10.1109/TNNLS.2022.3229161.
3. Grinsztajn L., Oyallon E., Varoquaux G. Why do tree-based models still outperform deep learning on typical tabular data? Adv. Neural Inf. Process. Syst. 2022. Vol. 35. Р. 507–520.
4. Qiu Q., Liu H. Numerical Embedding of Categorical Features in Tabular Data: A Survey. 2023 International Conference on Machine Learning and Cybernetics (ICMLC). Adelaide, Australia: IEEE, 2023. P. 446–451. DOI: https://doi.org/10.1109/ICMLC58545.2023.10327921.
5. Breiman L. Random Forests. Mach. Learn. 2001. Vol. 45, N 1. Р. 5–32. DOI: https://doi.org/10.1023/A:1010933404324.
6. Chen T., Guestrin C. XGBoost: A Scalable Tree Boosting System. 2016. Jun. 10, 2016. arXiv: arXiv:1603.02754. DOI: https://doi.org/10.48550/arXiv.1603.02754.
7. Hutter F., Kotthoff L., Vanschoren J. Automated Machine Learning: Methods, Systems, Challenges. The Springer Series on Challenges in Machine Learning. Cham: Springer International Publishing, 2019. DOI: https://doi.org/10.1007/978-3-030-05318-5.
8. Голуб С.В. Багаторівневе моделювання в технологіях моніторингу оточуючого середовища: монографія. Черкаси: Вид. від. ЧНУ імені Богдана Хмельницького, 2007. 220 с. ISBN 978-966-353-062-8.
9. Akiba T., Sano S., Yanase T., Ohta T., Koyama M. Optuna: A Next-generation Hyperparameter Optimization Framework. 2019. Jul. 25. arXiv: arXiv:1907.10902. URL: http://arxiv.org/abs/1907.10902 (аccessed: 10.11.2024).
10. Watanabe S. Tree-Structured Parzen Estimator: Understanding Its Algorithm Components and Their Roles for Better Empirical Performance. 2023. May 2. arXiv: arXiv:2304.11127. URL: http://arxiv.org/abs/2304.11127 (аccessed: 10.11.2024).
11. Hansen N. The CMA Evolution Strategy: A Tutorial. 2023. Mar. 10. arXiv: arXiv:1604.00772. URL: http://arxiv.org/abs/1604.00772 (аccessed: 10.11.2024).
12. Non-Dominated Sorting Genetic Algorithm II - an overview | ScienceDirect Topics. URL: https://www.sciencedirect.com/topics/engineering/non-dominated-sorting-genetic-algorithm-ii (аccessed: 10.11.2024).
13. Колос П.О., Голуб С.В. Умови конструювання алгоритмів синтезу моделей у системах багаторівневого перетворення інформації. Вісник Східноукраїнського національного університету імені Володимира Даля. 2009. № 6 (136), Ч. 1. С. 325–329.
14. Голуб С.В., Колос П.О. Застосування стратегії оптимальності при виборі алгоритмів синтезу моделей у системах багаторівневого соціоекологічного моніторингу. Математичні машини і системи. 2010. № 4. С. 127–134.
15. Farris F.A. The Gini Index and Measures of Inequality. Am. Math. Mon. 2010. Vol. 117, N 10. P. 851–864. DOI: https://doi.org/10.4169/000298910x523344.
16. Geurts P., Ernst D., Wehenkel L. Extremely randomized trees. Mach. Learn. 2006. Vol. 63, N 1. P. 3–42. DOI: https://doi.org/10.1007/s10994-006-6226-1.
17. Lundberg S., Lee S.-I. A Unified Approach to Interpreting Model Predictions. 2017. Nov. 25. arXiv: arXiv:1705.07874. DOI: https://doi.org/10.48550/arXiv.1705.07874.
18. Zhou Z.-H., Feng J. Deep forest. Natl. Sci. Rev. Vol. 6, N 1. P. 74–86. DOI: https://doi.org/10.1093/nsr/nwy108.
19. Ke G. et al. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. Advances in Neural Information Processing Systems, Curran Associates, Inc., 2017. URL: https://proceedings.neurips.cc/paper_files/paper/2017/hash/6449f44a102fde848669bdd9eb6b76faAbstract.html (аccessed: 09.11.2024).
20. Prokhorenkova L., Gusev G., Vorobev A., Dorogush A.V., Gulin A. CatBoost: unbiased boosting with categorical features. 2019. Jan. 20. arXiv: arXiv:1706.09516. DOI: https://doi.org/10.48550/arXiv.1706.09516.
21. Duan T. et al. NGBoost: Natural Gradient Boosting for Probabilistic Prediction. 2020. Jun. 09. arXiv: arXiv:1910.03225. DOI: https://doi.org/10.48550/arXiv.1910.03225.
22. Zhang H., Si S., Hsieh C.-J. GPU-acceleration for Large-scale Tree Boosting. 2017. Jun. 26. arXiv: arXiv:1706.08359. DOI: https://doi.org/10.48550/arXiv.1706.08359. ISSN 1028-9763. Математичні машини і системи. 2025. № 2 95
23. Arik S.O., Pfister T. TabNet: Attentive Interpretable Tabular Learning. 2020. Dec. 09. arXiv: arXiv:1908.07442. URL: http://arxiv.org/abs/1908.07442 (аccessed: 10.11.2024).
24. Luo J., Xu S. NCART: Neural Classification and Regression Tree for Tabular Data. 2024. Feb. 28.
arXiv: arXiv:2307.12198. URL: http://arxiv.org/abs/2307.12198 (аccessed: 10.11.2024).
25. Yang Y., Morillo I.G., Hospedales T.M. Deep Neural Decision Trees. 2018. Jun. 19. arXiv: arXiv:1806.06988. URL: http://arxiv.org/abs/1806.06988 (аccessed: 10.11.2024).
26. Katzir L., Elidan G., El-Yaniv R. Net-DNF: Effective Deep Modeling of Tabular Data. International Conference on Learning Representations. 2024. Nov. 10. URL: https://openreview.net/forum?id=73WTGs96kho (аccessed: 10.11.2024).
27. Hastie T., Tibshirani R., Wainwright M. Statistical Learning with Sparsity: The Lasso and Generalizations. Chapman & Hall/CRC, 2015. P. 7–23.
28. Hoerl A.E., Kennard R.W. Ridge Regression: Biased Estimation for Nonorthogonal Problems. Technometrics. 2000. Vol. 42, N 1. P. 80–86. DOI: https://doi.org/10.2307/1271436.
29. Tibshirani R. Regression Shrinkage and Selection via the Lasso. J. R. Stat. Soc. Ser. B Methodol. 1996. Vol. 58, N 1. P. 267–288.
30. Zou H., Hastie T. Regularization and Variable Selection Via the Elastic Net. J. R. Stat. Soc. Ser. B Stat. Methodol. 2005. Vol. 67, N 2. P. 301–320. DOI: https://doi.org/10.1111/j.14679868.2005.00503.x.
31. Pattaro E.U. iFood Marketing Analytics. 2020. URL: https://github.com/nailson/ifood-data-businessanalyst-test.
32. Taxi Price Prediction. 2024. URL: https://www.kaggle.com/datasets/denkuznetz/taxi-priceprediction/data.
33. Beridze G. Diamond Online Marketplace. 2024. URL: https://www.kaggle.com/datasets/beridzeg45/diamonds-prices-prediction/data.

