<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE root>
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" article-type="research-article" dtd-version="1.2" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">News of the Kabardino-Balkarian Scientific Center of the Russian Academy of Sciences</journal-id><journal-title-group><journal-title xml:lang="en">News of the Kabardino-Balkarian Scientific Center of the Russian Academy of Sciences</journal-title><trans-title-group xml:lang="ru"><trans-title>Известия Кабардино-Балкарского научного центра РАН</trans-title></trans-title-group></journal-title-group><issn publication-format="print">1991-6639</issn><issn publication-format="electronic">2949-1940</issn></journal-meta><article-meta><article-id pub-id-type="publisher-id">290710</article-id><article-id pub-id-type="doi">10.35330/1991-6639-2025-27-1-143-151</article-id><article-id pub-id-type="edn">XRYMDH</article-id><article-categories><subj-group subj-group-type="toc-heading" xml:lang="en"><subject>System analysis, management and information processing</subject></subj-group><subj-group subj-group-type="toc-heading" xml:lang="ru"><subject>Системный анализ, управление и обработка информации</subject></subj-group><subj-group subj-group-type="article-type"><subject>Research Article</subject></subj-group></article-categories><title-group><article-title xml:lang="en">Comparative analysis of class imbalance reduction methods in building machine learning models in the financial sector</article-title><trans-title-group xml:lang="ru"><trans-title>Сравнительный анализ методов снижения дисбаланса классов при построении моделей машинного обучения в финансовом секторе</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0000-9591-3301</contrib-id><contrib-id contrib-id-type="spin">3088-3121</contrib-id><name-alternatives><name xml:lang="en"><surname>Konstantinov</surname><given-names>Alexey F.</given-names></name><name xml:lang="ru"><surname>Константинов</surname><given-names>Алексей Федорович</given-names></name></name-alternatives><address><country country="RU">Russian Federation</country></address><bio xml:lang="ru"><p>аспирант кафедры информатики</p></bio><bio xml:lang="en"><p>Post-graduate Student, Department of Informatics</p></bio><email>konstantinovaf@gmail.com</email><xref ref-type="aff" rid="aff1"/></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-5229-8070</contrib-id><contrib-id contrib-id-type="spin">2513-8831</contrib-id><name-alternatives><name xml:lang="ru"><surname>Дьяконова</surname><given-names>Людмила Павловна</given-names></name><name xml:lang="en"><surname>Dyakonova</surname><given-names>Lyudmila P.</given-names></name></name-alternatives><address><country country="RU">Russian Federation</country></address><bio xml:lang="ru"><p>канд. физ.-мат. наук, доцент, кафедра информатики</p></bio><bio xml:lang="en"><p>Candidate of Physical and Mathematical Sciences, Associate Professor, Department of Informatics</p></bio><email>Dyakonova.LP@rea.ru</email><xref ref-type="aff" rid="aff1"/></contrib></contrib-group><aff-alternatives id="aff1"><aff><institution xml:lang="ru">Российский экономический университет им. Г. В. Плеханова</institution></aff><aff><institution xml:lang="en">Plekhanov Russian University of Economics</institution></aff></aff-alternatives><content-language>ru</content-language><pub-date date-type="pub" iso-8601-date="2025-02-15" publication-format="electronic"><day>15</day><month>02</month><year>2025</year></pub-date><pub-date date-type="collection"><year>2025</year></pub-date><volume>27</volume><issue>1</issue><issue-title xml:lang="ru"/><issue-title xml:lang="en"/><fpage>143</fpage><lpage>150</lpage><history><date date-type="received" iso-8601-date="2025-05-07"><day>07</day><month>05</month><year>2025</year></date><date date-type="accepted" iso-8601-date="2025-05-07"><day>07</day><month>05</month><year>2025</year></date></history><permissions><copyright-statement xml:lang="en">Copyright ©; 2025, Константинов А.F., Дьяконова Л.P.</copyright-statement><copyright-statement xml:lang="ru">Copyright ©; 2025, Константинов А.Ф., Дьяконова Л.П.</copyright-statement><copyright-year>2025</copyright-year><copyright-holder xml:lang="en">Константинов А.F., Дьяконова Л.P.</copyright-holder><copyright-holder xml:lang="ru">Константинов А.Ф., Дьяконова Л.П.</copyright-holder><ali:free_to_read xmlns:ali="http://www.niso.org/schemas/ali/1.0/"/><license><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/">https://creativecommons.org/licenses/by/4.0</ali:license_ref></license></permissions><self-uri xlink:href="https://journals.rcsi.science/1991-6639/article/view/290710">https://journals.rcsi.science/1991-6639/article/view/290710</self-uri><abstract xml:lang="en"><p>The article discusses methods for improving quality metrics of machine learning models used in the financial sector. Due to the fact that the data sets on which the models are trained have class imbalances, it is proposed to use models aimed at reducing the imbalance. The study conducted experiments using 9 methods for accounting for class imbalances with three data sets on retail lending. The CatboostClassifier gradient boosting model, which does not take into account class imbalances, was used as the base model. The experiments showed that the use of the RandomOverSampler method provides a significant increase in classification quality metrics compared to the base model. The results indicate the promise of further research into methods for accounting for class imbalances in the study of financial data, as well as the feasibility of application of the considered methods in practice.</p></abstract><trans-abstract xml:lang="ru"><p>В статье рассматриваются методы улучшения показателей качества моделей машинного обучения, применяемых в финансовом секторе. В связи с тем, что наборы данных, на которых обучаются модели, обладают несбалансированностью классов, предлагается использовать модели, направленные на снижение дисбаланса. В исследовании были проведены эксперименты с применением 9 методов учета несбалансированности классов к трем наборам данных по розничному кредитованию. В качестве базовой использовалась модель градиентного бустинга CatboostClassifier, не учитывающая дисбаланс классов. Проведенные эксперименты показали, что применение метода RandomOverSampler дает существенный прирост показателей качества классификации по сравнению с базовой моделью. Результаты свидетельствуют о перспективности дальнейших исследований методов учета дисбаланса классов при изучении финансовых данных, а также о целесообразности применения рассмотренных методов на практике.</p></trans-abstract><kwd-group xml:lang="en"><kwd>financial risks</kwd><kwd>machine learning</kwd><kwd>classification</kwd><kwd>class imbalance</kwd></kwd-group><kwd-group xml:lang="ru"><kwd>финансовые риски</kwd><kwd>машинное обучение</kwd><kwd>классификация</kwd><kwd>дисбаланс классов</kwd></kwd-group><funding-group/></article-meta></front><body></body><back><ref-list><ref id="B1"><label>1.</label><mixed-citation>Chawla N.V., Bowyer K.W., Hall L.O., Kegelmeyer W.P. Smote: synthetic minority over-sampling technique. Journal of artificial intelligence research. 2002. Vol. 16. Pp. 321–357. DOI: 10.1613/jair.953</mixed-citation></ref><ref id="B2"><label>2.</label><mixed-citation>He H., Bai Y., Garcia E.A., Li S. Adasyn: adaptive synthetic sampling approach for imbalanced learning. In 2008 IEEE International Joint Conference on Neural Networks (IEEE World Congress on Computational Intelligence). 2008. Pp. 1322–1328. DOI: 10.1109/IJCNN.2008.4633969</mixed-citation></ref><ref id="B3"><label>3.</label><mixed-citation>Han H., Wang W.-Y., Mao B.-H. Borderline-smote: a new over-sampling method in imbalanced data sets learning. International conference on intelligent computing. 2005. Pp. 878–887. Springer. DOI: 10.1007/11538059_91</mixed-citation></ref><ref id="B4"><label>4.</label><mixed-citation>Tomek I. Two modifications of cnn. IEEE Trans. Systems, Man and Cybernetics. 1976. Vol. 6. Pp. 769–772. DOI: 10.1109/TSMC.1976.4309452</mixed-citation></ref><ref id="B5"><label>5.</label><mixed-citation>Laurikkala J. Improving identification of difficult small classes by balancing class distribution. In Conference on Artificial Intelligence in Medicine in Europe. 2001. Pp. 63–66. Springer. DOI: 10.1007/3-540-48229-6_9</mixed-citation></ref><ref id="B6"><label>6.</label><mixed-citation>Batista G., Prati R.C., Monard M.C. A study of the behavior of several methods for balancing machine learning training data. ACM Sigkdd Explorations Newsletter 2004. Vol. 6. No. 1. Pp. 20–29. DOI: 10.1145/1007730.1007735</mixed-citation></ref><ref id="B7"><label>7.</label><mixed-citation>Batista G., Bazzan B., Monard M., Balancing Training Data for Automated Annotation of Keywords: a Case Study. In WOB. 2003. Pp. 10–18. BibTeX key: conf/wob/BatistaBM03</mixed-citation></ref></ref-list></back></article>
