Analisis Perbandingan Imputasi Mean, Median, dan KNN pada Decision Tree Berbasis Entropi untuk Penanganan Nilai Hilang dalam Prediksi Komplikasi Infark Miokard
DOI:
https://doi.org/10.38035/dit.v4i1.3595Keywords:
Data Imputasi, Decision Tree Berbasis Entropi, Infark Miokard, KNN Imputer, Missing ValuesAbstract
Dataset klinis sering memuat nilai hilang yang dapat memengaruhi keandalan model prediksi. Penelitian ini membandingkan imputasi mean, median, dan KNN dalam menangani nilai hilang pada dataset Myocardial Infarction Complications. Data awal terdiri atas 1.700 rekam pasien, 111 fitur masukan, dan LET_IS sebagai target prediksi. Setelah simbol nilai hilang dikonversi, kolom dengan nilai hilang tinggi dibersihkan, dan baris yang tidak dapat digunakan pada target penelitian dihapus, data yang diproses pada tahap train-test split menjadi 1.350 rekam, yaitu 1.080 data latih dan 270 data uji. Setiap skenario imputasi dievaluasi menggunakan Decision Tree berbasis Entropi sebagai pendekatan praktis terhadap prinsip pemisahan C4.5. Hasil Google Colab terbaru pada skenario single-split menunjukkan bahwa imputasi median menghasilkan akurasi tertinggi sebesar 89,63%, diikuti KNN sebesar 88,52% dan mean sebesar 88,15%. Rekomendasi utama penelitian ini adalah median karena memberikan akurasi tertinggi pada pengujian utama.
References
Afkanpour, M., Hosseinzadeh, E., & Tabesh, H. (2024). Identify the Most Appropriate Imputation Method for Handling Missing Values in Clinical Structured Datasets: A Systematic Review. BMC Medical Research Methodology, 24(1), 188. https://doi.org/10.1186/s12874-024-02310-6
Breiman, L., Friedman, J. H., Olshen, R. A., & Stone, C. J. (1984). Classification and Regression Trees. Wadsworth.
Ghafari, R., & others. (2023). Prediction of the Fatal Acute Complications of Myocardial Infarction via Machine Learning Algorithms. Journal of Tehran University Heart Center, 18(4), 278–287. https://doi.org/10.18502/jthc.v18i4.14827
Golovenkin, S. E., & others. (2020). Trajectories, Bifurcations, and Pseudo-Time in Large Clinical Datasets: Applications to Myocardial Infarction and Diabetes Data. GigaScience, 9(11), giaa128. https://doi.org/10.1093/gigascience/giaa128
Harris, C. R., & others. (2020). Array Programming with NumPy. Nature, 585, 357–362. https://doi.org/10.1038/s41586-020-2649-2
Khamis, G. S. M., Mohammed, Z. M. S., Alanazi, S. M., Mahmoud, A. F. A., Abdalla, F. A., & Bkheet, S. A. (2024). Prediction of Myocardial Infarction Complications Using Gradient Boosting. Engineering, Technology & Applied Science Research, 14(5), 17486–17492.
Li, J., & others. (2024). Comparison of the Effects of Imputation Methods for Missing Data in Predictive Modelling of Cohort Study Datasets. BMC Medical Research Methodology, 24, 41. https://doi.org/10.1186/s12874-024-02173-x
Little, R. J. A., & Rubin, D. B. (2019). Statistical Analysis with Missing Data (3rd ed.). Wiley.
Liu, M., & others. (2023). Handling Missing Values in Healthcare Data: A Systematic Review of Deep Learning-Based Imputation Techniques. Artificial Intelligence in Medicine, 142, 102587. https://doi.org/10.1016/j.artmed.2023.102587
Luque, A., Carrasco, A., Martin, A., & de Las Heras, A. (2019). The Impact of Class Imbalance in Classification Performance Metrics Based on the Binary Confusion Matrix. Pattern Recognition, 91, 216–231. https://doi.org/10.1016/j.patcog.2019.02.023
Mazdadi, M. I., Saragih, T. H., Budiman, I., Farmadi, A., & Tajali, A. (2024). The Effectiveness of Data Imputations on Myocardial Infarction Complication Classification Using Machine Learning Approach with Hyperparameter Tuning. Jurnal Ilmiah Teknik Elektro Komputer Dan Informatika, 10(3), 520–533. https://doi.org/10.26555/jiteki.v10i3.29479
Pedregosa, F., & others. (2011). Scikit-Learn: Machine Learning in Python. Journal of Machine Learning Research, 12, 2825–2830.
Powers, D. M. W. (2011). Evaluation: From Precision, Recall and F-Measure to ROC, Informedness, Markedness and Correlation. Journal of Machine Learning Technologies, 2(1), 37–63.
Quinlan, J. R. (1993). C4.5: Programs for Machine Learning. Morgan Kaufmann.
Saito, T., & Rehmsmeier, M. (2015). The Precision-Recall Plot is More Informative Than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLoS ONE, 10(3), e0118432. https://doi.org/10.1371/journal.pone.0118432
Satty, A. A., Salih, M. M. Y., Hassaballa, A. A., Gumma, E. A. E., Abdallah, A., & Khamis, G. S. M. (2024). Comparative Analysis of Machine Learning Algorithms for Investigating Myocardial Infarction Complications. Engineering, Technology & Applied Science Research, 14(1), 12775–12779. https://doi.org/10.48084/etasr.6691
Thygesen, K., & others. (2018). Fourth Universal Definition of Myocardial Infarction (2018). Circulation, 138(20), e618–e651. https://doi.org/10.1161/CIR.0000000000000617
van Buuren, S. (2018). Flexible Imputation of Missing Data (2nd ed.). CRC Press.
.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Muhammad Rizal Arifin, Zainal Abidin

This work is licensed under a Creative Commons Attribution 4.0 International License.
Copyright :
Authors who publish their manuscripts in this journal agree to the following conditions:
- Copyright in each article belongs to the author.
- The author acknowledges that the DIT has the right to be the first to publish under a Creative Commons Attribution 4.0 International license (Attribution 4.0 International CC BY 4.0).
- Authors can submit articles separately, arrange the non-exclusive distribution of manuscripts that have been published in this journal to other versions (for example, sent to the author's institutional repository, publication in a book, etc.), by acknowledging that the manuscript has been published for the first time at DIT.




















