Comparative Analysis of Machine Learning Models For Predicting Default in Home Credit Companies
Downloads
Credit scoring is an important process in the financial world to accurately identify the risk of prospective debtors. This study aims to build a credit scoring model by comparing the performance of several machine learning algorithms, namely XGBoost, random forest, and logistic regression. The dataset used is part of the training data and test data with several ratios, namely 70:30, 75:25, and 80:20. The preparation process is carried out through variable selection using the 5C principle, missing value imputation, categorical transformation of variables, and creation of derived features. Furthermore, modeling and optimization are carried out for each model to improve classification performance, especially in recognizing debtors who have the potential to default. The evaluation results show that the XGBoost model has the best performance with an accuracy of 84.8%, a precision of 84.9% and a recall of 84.7%, and an AUC of 92.3%. The main assessment of the character principle is the external credit score variable, the main assessment of the capacity principle is income, the main assessment of the capital principle is car ownership, the main assessment of the collateral principle is the credit financing ratio, and the main assessment of the condition principle is the regional rating.
Abdallah, Z. S., & Webb, G. (2017). Encyclopedia of Machine Learning and Data Mining. Encycl. Mach. Learn. Data Min.
Abdullah, T. (2012). Bank dan Lembaga Keuangan. Jakarta: PT. Raja Grafindo Persada.
Akinjole, A., Shobayo, O., Popoola, J., Okoyeigbo, O., & Ogunleye, B. (2024). Ensemble-Based Machine Learning Algorithm for Loan Default Risk Prediction. Multidisciplinary Digital Publishing Institute, 3423.
Ali, J., Khan, R., Ahmad, N., & Maqsood, I. (2012). Random Forests and Decision Trees. International Journal of Computer Science, 272-278.
Andriani, W., Gunawan, G., & Naja, N. N. (2025). Analisis Perbandingan Machine Learning untuk Prediksi Kelayakan Kredit Perbankan pada Bank BRI Tegal. Jurnal Penerapan Teknologi Informasi dan Komunikasi, 82-92.
Awangga, R. M., & Khonsa, N. H. (2022). Analisis Performa Algoritma Random Forest dan Naive Bayes Multinomial pada Dataset Ulasan Obat dan Ulasan Film. Jurnal Telekomunikasi dan Komputer, 60–70.
Brigham, E., & Houston, J. (2019). Fundamentals of Financial Management 15e. Boston: Cengage Learning.
Budi, E. S., Chan, A. N., Alda, P. P., & Idris, M. A. (2024). Optimasi Model Machine Learning untuk Klasifikasi dan Prediksi Citra. Rekayasa Teknik Informatika dan Informasi, 502-509.
Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic Minority Over-Sampling Technique. Journal of Artificial Intelligence Research, 321-357.
Chen, T., & Guestrin, C. (2016). Xgboost: A Scalable Tree Boosting. In Proceedings Of The 22nd ACM International Conference On Knowledge Discovery And Data, 785–794.
Creswell, J. W. (2014). Research Design: Qualitative, Quantitative, and Mixed Methods Approaches. London: SAGE Publications.
Devi, & Raju. (2024). A Comparative Analysis of Machine Learning Algorithms for Big Data. International Journal of Management Science and Engineering Management, 1608–1630.
Dr. Zainuddin Iba, S. M. (2023). Metode Penelitian. Jawa Tengah: Eureka Media Aksara.
Fathurohman, A. (2021). Machine Learning Untuk Pendidikan: Mengapa dan Bagaimana. Jurnal Informatika dan Teknologi Komputer, 57-62.
Fitria, E. R., & Rozci, F. (2022). Penerapan Metode Regresi Least Absolute Shrinkge And Selection Operator (LASSO) Dan Regresi Linier Untuk Prediksi Tingkat Kemiskinan Di Indonesia. Jurnal Ilmiah Sosio Agribis, 123-132.
Hamonangan. (2020). Analisis Penerapan Prinsip 5C Dalam Penyaluran Pembiayaan Pada Bank Muamalat KCU Padangsidempuan. Jurnal Ilmiah MEA (Manajemen, Ekonomi, dan Akuntansi), 454-466.
Hanafi. (2023). Data Cleaning dalam Big Data : Review. Jakarta: ResearchGate.
Hand, D. J., & Henley, W. E. (1997). Statistical Classification Methods in Consumer Credit Scoring: A Review. Journal of Royal Statistical Social, 523-541.
Hosmer, D. W., Lemeshow, S., & Sturdivant, R. X. (2013). Applied Logistic. New Jersey: John Wiley & Sons, Inc.
Jacobs, M. (2010). An Empirical Study of Exposure at Default. Journal of Advanced Studies in Finance, 31-59.
Jiajia, F., Jiangfeng, X., & Junfeng, Z. (2021). Intrusion Detection Model Based on SAE BALSTM. 2021 IEEE International Conference on Artificial Intelligence and Computer, 1192–1197.
Jolliffe, I. (2002). Principal Component Analysis. Heidelberg: Springer.
Kasmir. (2016). Analisis Laporan Keuangan. Jakarta: RajaGrafindo Persada.
Kornfeld, S. (2020). Predicting Default Probability in Credit Risk using Machine Learning Algorithms. Royal Institute of Technology School of Engineering Sciences, 186.
Kurniasih, D., Rusfiana, Y., Subagyo, A., & Nuradhawati, R. (2021). Teknik Analisis Data. Bandung: Alfabeta.
Lestari, K. S., & Sudiana, I. K. (2019). Pengaruh Lama Kerja, Umur, Dan Tingkat Pendidikan Terhadap Produktivitas Dan Pendapatan. E-Jurnal Ekonomi Pembangunan Universitas Udayana, 1575-1607.
Lusiyanti, D., Musdalifah, S., Sahari, A., & Al Fajri, I. (2025). Evaluasi Kinerja Algoritma Machine Learning. Junal Matematika, 84-92.
Mahandiri, D. N., & Muliantara, A. (2024). Default Risk Prediction Using Decision Tree Study Case of Home Credit. Jurnal Elektronik Ilmu Komputer Udayana, 647-654.
Mann, H. B., & Whitney, D. R. (1947). On a Test of Whether One of Two Random Variables is Stochastically Larger than The Other. United States: Annals of Mathematical Statistics.
McHugh, M. L. (2013). The Chi-square test of independence. Biochemia Medica, 143–149.
Mustafa, & Cudi, M. (2023). A Comprehensive Review of Feature Selection and Feature Selection Stability. Journal Of Science, 1506-1520.
Nasution, D. A., Khotimah, H. H., & Chamidah, N. (2019). Perbandingan Normalisasi Data Untuk Klasifikasi Wine Menggunakan Algoritma K-NN. Journal of Computer Engineering System and Science, 78-82.
Nurhopipah, A., & Hasanah, U. (2020). Dataset Splitting Techniques Comparison For Face. Indonesian J. Comput. Cybern. Syst, 341–352.
Omarzai, F. (2024, July 21). XGBoost Classification In Depth. p. 1.
Pachamanova, D., & Fabozzi, F. J. (2010). Simulation and Optimization Modeling in Finance. Hoboken: John Wiley & Son.
Paul, A., Mukherjee, D. P., Das, P., Gangopadhyay, A., Chintha, A. R., & Kundu, S. (2018). Improved Random Forest for Classification. IEEE Transactions on Image Processing, 4012–4024.
Priliyanabita, G., Irawan, A., & Putri, E. S. (2024). Pengaruh Tingkat Pendidikan, Jenis Pekerjaan, dan Tingkat Pendapatan Terhadap Pembayaran Pajak Bumi dan Bangunan. Jurnal Pendidikan Tambusai, 5222-5234.
Redman, T. C. (1996). Data Quality for the Information Age. Norwood: Artech House.
Rianto, H., & Wahono, R. S. (2015). Resampling Logistic Regression untuk Penanganan Ketidakseimbangan Class pada Prediksi Cacat Software. J. Softw. Eng., 46–53.
Rodríguez, R. A. (2024). A Required Debt Service Coverage Ratio Related to the Economic Value of the Asset Involved. Journal of Financial Risk Management, 618-642.
Roihan, A., Abas Sunarya, P., & Rafika, A. (2019). Pemanfaatan Machine Learning dalam Berbagai Bidang: Review paper. Indonesian Journal on Computer and Information Technology, 75-82.
Salter, R. (2023). Explainable Artificial Intelligence and its Apllications in Behavioral Credit Scoring. Digitala Vetenskapliga Arkivet, 1784385.
Santoso, R., Megasar, R., & Hambali, Y. (2020). Implementasi Metode Machine Learning Menggunakan Algoritma Evolving Artificial Neural Network Pada Kasus Prediksi Diagnosis Diabetes Implementation of Machine Learning Method Using Evolving Artificial Neural Network Algorithm in Prediction of Diabetes. Jurnal Aplikasi dan Teori Ilmu Komputer, 85-97.
Soman, K., Loganathan, R., & Ajay, V. (2009). Machine Learning with SVM and Other Kernel Methods. PHI Learning Pvt. Ltd.
Sugianto, Widyasari, Y. D., & Wardhani, K. D. (2024). Modeling and Application of Credit Scoring Based on A Multi-Objective Approach to Debtor Data in PT. Bank Riau Kepri. International Journal On Informatics Visualization, 220-230.
Suhadolnik, N., Ueyama, J., & Silva, S. D. (2023). Machine Learning for Enhanced Credit Risk Assessment: An Ampirical Approach. Journal Of Risk And Financial Management, 496.
Sukmaningrum, D. A. (2023). Analisa Kelayakan Nasabah Menggunakan Metode Prinsip 5c Dalam Pembiayaan KPR. Jurnal Ekonomi Manajemen dan Sosial, 32-42.
Sutoyo, S. (2002). Manajemen Keuangan Bagi Eksekutif Non Keuangan. Jakarta: Damar Mulia Pustaka.
Thavichaigarn, N. (2022). Corporate Credit Rating Prediction Using Deep Learning. Chulalongkorn University Theses And Dissertations, 5808.
Varian, H. R. (2014). Big Data: New Tricks for Econometrics. Journal of Economic Perspectives, 3-28.
Walpole, R. E. (1995). Pengantar Statistika. Jakarta: PT. Gramedia Pustaka Utama.
Wang, Y., Zhang, Y., Lu, Y., & Yu, X. (2020). A Comparative Assessment of Credit Risk Model Based on Machine Learning. Procedia Computer Science, 141-149.
Weil, R. L., Schipper, K., & Francis, J. (2014). Financial accounting. Fourteenth edition. South-Western: Ohio.
Zhou, Y. (2022). Loan Default Prediction Based on Machine Learning Methods. European Union Digital Library, 2328740.
Copyright (c) 2025 Daniel Alexander, Raden Supriyanto

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution-ShareAlike 4.0 International. that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.







