Small-Sample Classification of Economic Impacts from Registered SMEs Using Machine Learning
Main Article Content
Abstract
Registered Small and Medium Enterprises (SMEs) serve as a fundamental pillar of Thailand’s economy; however, modeling their economic impact is often constrained by the limitations of small-scale datasets. This study evaluates the performance of seven machine learning algorithms—Decision Tree, Support Vector Machine (SVM), Gradient Boosting, K-Nearest Neighbor (KNN), Naïve Bayes, XGBoost, and Multi-Layer Perceptron (MLP) Neural Network—in classifying the economic contributions of registered SMEs. To mitigate issues arising from class imbalance and limited sample size, the SMOTEENN (Synthetic Minority Over-sampling Technique and Edited Nearest Neighbors) hybrid sampling method was employed. The results indicate that while SMOTEENN significantly improved recall (reaching 1.000 for the majority of models), varying precision levels resulted in moderate F1 scores for the Decision Tree (0.83), SVM (0.87), and Gradient Boosting (0.88) classifiers. Notably, XGBoost and Naïve Bayes achieved peak performance with perfect scores (1.000) across accuracy, F1 score, and AUC metrics, demonstrating a robust capacity for pattern recognition within small-sample contexts, albeit with a recognized risk of overfitting. In contrast, the MLP Neural Network exhibited the lowest efficacy (Accuracy 0.778, AUC 0.438), suggesting that deep learning architectures may require more extensive hyperparameter tuning or larger datasets to remain competitive. These findings suggest that ensemble methods and probabilistic models are most effective for SME classification tasks under data constraints, providing a framework for more precise, data-driven economic policy-making.
Article Details
References
- R. Rodhiyah, Dampak Sosial Ekonomi Keberadaan Usaha Kecil Menengah (UKM) Konveksi Di Kota Semarang, J. ILMU Sos. 14 (2016), 1–14. https://doi.org/10.14710/jis.14.1.2015.1-14.
- C.L. Cofino, P.D. Cerna, SME-ODA: A Machine Learning Driven Platform for Data Integration and Analysis in Small and Medium Enterprises (SMEs), Int. J. Comput. Sci. Mob. Comput. 14 (2025), 67–78. https://doi.org/10.47760/ijcsmc.2025.v14i06.008.
- J.C. Fernandez de Arroyabe, M.F. Arroyabe, I. Fernandez, C.F.A. Arranz, Cybersecurity Resilience in SMEs. A Machine Learning Approach, J. Comput. Inf. Syst. 64 (2023), 711–727. https://doi.org/10.1080/08874417.2023.2248925.
- Inayatulloh, Machine Learning Adoption Model for SME E-Commerce Enhancement, in: 2023 IEEE 3rd International Maghreb Meeting of the Conference on Sciences and Techniques of Automatic Control and Computer Engineering (MI-STA), IEEE, 2023, pp. 281–284. https://doi.org/10.1109/mi-sta57575.2023.10169335.
- J. Jiang, C. Zhang, L. Ke, N. Hayes, Y. Zhu, et al., A Review of Machine Learning Methods for Imbalanced Data Challenges in Chemistry, Chem. Sci. 16 (2025), 7637–7658. https://doi.org/10.1039/d5sc00270b.
- P. Pangestu, R. Novita, M. Mustakim, Systematic Literature Review: Perbandingan Algoritma Klasifikasi, INOVTEK Polbeng - Seri Inform. 8 (2023), 431. https://doi.org/10.35314/isi.v8i2.3698.
- S. Ghildiyal, R. Garg, S. Dhanik, A. Sarkar, G. Sharma, et al., A Comprehensive Review Analysis of Supervised Machine Learning Techniques, in: 2024 5th International Conference on Smart Electronics and Communication (ICOSEC), IEEE, 2024, pp. 925–930. https://doi.org/10.1109/icosec61587.2024.10722516.
- M.J. Setiawan, V.R.S. Nastiti, DANA App Sentiment Analysis: Comparison of XGBoost, SVM, and Extra Trees, J. Sisfokom Sist. Inf. Komput. 13 (2024), 337–345. https://doi.org/10.32736/sisfokom.v13i3.2239.
- I.R. Hendrawan, E. Utami, A.D. Hartanto, Comparison of Naïve Bayes Algorithm and XGBoost on Local Product Review Text Classification, Edumatic: J. Pendidik. Inform. 6 (2022), 143–149. https://doi.org/10.29408/edumatic.v6i1.5613.
- N.V. Mmbengeni, M.E. Mavhungu, M. John, The Effects of Small and Medium Enterprises on Economic Development in South Africa, Int. J. Entrep. 25 (2021), 1–11.
- S. Intharasompong, A. Jitpattanakul, S. Mekruksavanich, P. Wongwiwat, W. Phaphan, Classification Models and Variable Selection for SME Credit Risk Assessment in Balance the Data Set, in: 2025 Joint International Conference on Digital Arts, Media and Technology with ECTI Northern Section Conference on Electrical, Electronics, Computer and Telecommunications Engineering (ECTI DAMT & NCON), IEEE, Nan, Thailand, 2025: pp. 493–498. https://doi.org/10.1109/ECTIDAMTNCON64748.2025.10961980.
- A. Bitetto, P. Cerchiello, S. Filomeni, A. Tanda, B. Tarantino, Machine Learning and Credit Risk: Empirical Evidence from Small- and Mid-Sized Businesses, Socioecon. Plan. Sci. 90 (2023), 101746. https://doi.org/10.1016/j.seps.2023.101746.
- S. Steinert, V. Ruf, D. Dzsotjan, N. Großmann, A. Schmidt, et al., A Refined Approach for Evaluating Small Datasets via Binary Classification Using Machine Learning, PLOS ONE 19 (2024), e0301276. https://doi.org/10.1371/journal.pone.0301276.
- I. Met, A. Erkoc, S.E. Seker, M.A. Erturk, B. Ulug, Product Recommendation System with Machine Learning Algorithms for SME Banking, Int. J. Intell. Syst. 2024 (2024), 5585575. https://doi.org/10.1155/2024/5585575.
- A. Vabalas, E. Gowen, E. Poliakoff, A.J. Casson, Machine Learning Algorithm Validation with a Limited Sample Size, PLOS ONE 14 (2019), e0224365. https://doi.org/10.1371/journal.pone.0224365.
- N.A.D. Suhaimi, H. Abas, A Systematic Literature Review on Supervised Machine Learning Algorithms, PERINTIS EJournal 10 (2021), 1–24.
- K. Kumar, B. Singh, Artificial Intelligence in the Startup World: A Bibliometric Study of Emerging Trends and Themes, Abhigyan 43 (2025), 315–337. https://doi.org/10.1177/09702385251349614.
- L. Marques, S. Moro, P. Ramos, Data-Driven Insights to Reduce Uncertainty from Disruptive Events in Passenger Railways, Public Transp. 17 (2025), 683–713. https://doi.org/10.1007/s12469-024-00380-9.
- K. Handayani, B. Lailiah, Comparison of XGboost, Extra Trees, and LightGBM with SMOTE for Fetal Health Classification, SISTEMASI 13 (2024), 980–990. https://doi.org/10.32520/stmsi.v13i3.3646.
- S. Rotbei, W.H. Tseng, B. Merino-Barbancho, M.S. Haleem, L. Montesinos, et al., Evaluating Impact of Movement on Diabetes via Artificial Intelligence and Smart Devices Systematic Literature Review, Expert Syst. Appl. 257 (2024), 125058. https://doi.org/10.1016/j.eswa.2024.125058.
- G.H. Sagala, D. Őri, Toward SMEs Digital Transformation Success: A Systematic Literature Review, Inf. Syst. E-Business Manag. 22 (2024), 667–719. https://doi.org/10.1007/s10257-024-00682-2.
- C. El Morr, M. Jammal, I. Bou-Hamad, S. Hijazi, D. Ayna, et al., Predictive Machine Learning Models for Assessing Lebanese University Students’ Depression, Anxiety, and Stress during COVID-19, J. Prim. Care Community Health 15 (2024), 21501319241235588. https://doi.org/10.1177/21501319241235588.
- Z. Xu, D. Shen, T. Nie, Y. Kou, A Hybrid Sampling Algorithm Combining M-Smote and ENN Based on Random Forest for Medical Imbalanced Data, J. Biomed. Inform. 107 (2020), 103465. https://doi.org/10.1016/j.jbi.2020.103465.
- E. Gürsoy, Y. Kaya, An Overview of Deep Learning Techniques for COVID-19 Detection: Methods, Challenges, and Future Works, Multimed. Syst. 29 (2023), 1603–1627. https://doi.org/10.1007/s00530-023-01083-0.
- H. He, E. Garcia, Learning from Imbalanced Data, IEEE Trans. Knowl. Data Eng. 21 (2009), 1263–1284. https://doi.org/10.1109/tkde.2008.239.
- N.V. Chawla, K.W. Bowyer, L.O. Hall, W.P. Kegelmeyer, SMOTE: Synthetic Minority Over-Sampling Technique, J. Artif. Intell. Res. 16 (2002), 321–357. https://doi.org/10.1613/jair.953.