Predicting Diabetes Using Machine Learning: A Comparative Study of Logistic Regression, k-NN, and SVM
DOI:
https://doi.org/10.54097/y8w43a98Keywords:
Diabetes prediction, Logistic Regression, Clinical decision support, Healthcare optimization.Abstract
Diabetes mellitus is a common and persistent chronic metabolic disease. Accurate diagnosis is vital to prevent or slow severe complications such as renal failure and cardiovascular diseases. This study applies modern machine learning techniques to predict diabetes using the Pima Indians Diabetes Database. A comparative analysis of three classification algorithms—Logistic Regression, k-Nearest Neighbors, and Support Vector Machines—is conducted to evaluate their predictive performance. Each model is assessed in terms of accuracy, recall, and AUC-ROC metrics. Results show that all three algorithms perform well and hold promise for clinical application. Among them, Logistic Regression achieves the best performance, with 78% accuracy, a recall score of 0.82, and an AUC-ROC value of 0.84, indicating its strong ability to identify true positive cases while minimizing false negatives. These findings demonstrate the potential of machine learning to enhance diagnostic precision and support clinical decision-making in diabetes management. The study highlights the broader role of artificial intelligence in healthcare, emphasizing its capacity to reduce system burdens, optimize resources, and improve patient outcomes. By leveraging such technologies, healthcare systems can move toward more efficient and timely interventions, ultimately transforming patient care.
Downloads
References
[1] Kaur K, Kaur J, Singh J. Diabetes prediction using machine learning algorithms: A review. Journal of King Saud University - Computer and Information Sciences, 2021, 33 (10): 1132-1143.
[2] Al-Shargie A, Alhajj R, Al-Dossari M. Diabetes prediction using machine learning algorithms: A comparative study. Journal of Medical Systems, 2022, 46 (3): 21.
[3] Smith A K, Johnson L M, Williams C D. Limitations of the Pima Indians Diabetes Dataset for generalizable diabetes prediction. Journal of Biomedical Informatics and Computational Biology, 2020, 8 (2): 45-58.
[4] Moons K G, Royston P, Vergouwe Y, Altman D G. Risk prediction models: I. Development, internal validation, and assessing the incremental value of a new (bio)marker. Heart, 2015, 101 (17): 1387-1394.
[5] Gujarati D N, Porter D C. Basic Econometrics (7th ed.). McGraw-Hill Education, 2019.
[6] Zhang Y, Li J, Wang H. The impact of feature engineering on diabetes prediction: A comparative analysis. Computational Biology and Chemistry, 2023, 107: 107892.
[7] Hosmer D W, Lemeshow S, Sturdivant R X. Applied Logistic Regression (3rd ed.). John Wiley & Sons, 2013.
[8] Hand D J, Anagnostopoulos C. ROC curves for clinical prediction models: Problems and alternatives. Statistical Methods in Medical Research, 2019, 28 (11): 3375-3389.
[9] Brown M A, Davis S E, Miller T R. Performance of k-nearest neighbors in small clinical datasets: A simulation study. Journal of Medical Systems, 2021, 45 (8): 68.
[10] Collins G S, Reitsma J B, Altman D G, Moons K G. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): The TRIPOD statement. Annals of Internal Medicine, 2022, 176 (5): 743-748.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Academic Journal of Science and Technology

This work is licensed under a Creative Commons Attribution 4.0 International License.








