Date of Award
2026-05-01
Degree Name
Master of Science
Department
Computational Science
Advisor(s)
Jonathon E. Mohl
Abstract
Prediabetes is a critical health condition that increases the risk of developing type 2 diabetes. The hemoglobin A1c (HbA1c) test diagnoses patients with prediabetes, but the disease has already caused metabolic alterations. Early detection is essential for timely interventions, and machine learning models offer a promising approach to identify prediabetic individuals through the analysis of biomarkers such as cytokines. We compared four classifiers-logistic regression, decision tree, random forest, and k-nearest neighbors - using cytokines (TNF-α, MCP-1, IL-1β, IL-6, IFN-γ), age, BMI, and waist-to-hip ratio (WHR). Models were evaluated using stratified 5-fold cross-validation and ROC-AUC. K-Nearest Neighbors (k-NN) achieved the highest performance (CV AUC = 0.771 ± 0.037), closely followed by random forest (0.750 ± 0.026), logistic regression (0.704 ± 0.027), and decision tree (0.637 ± 0.051). Feature importance varied by model: decision tree and random forest highlighted age, MCP-1, TNF-α, IL-6, and WHR as key predictors. IL-1β, IFN-γ, and BMI showed negligible importance. Random forest provided the best balance of predictive performance, sensitivity, and interpretability. These findings support the use of non-linear machine learning models for integrating clinical and biomarker data in chronic disease risk assessment.
Language
en
Provenance
Received from ProQuest
Copyright Date
2026-05
File Size
46 p.
File Format
application/pdf
Rights Holder
Luisa Veronica Gracia Mazuca
Recommended Citation
Gracia Mazuca, Luisa Veronica, "Prediabetes Prediction Before Disease Onset Using Multimodal Health Data: A Machine Learning Approach" (2026). Open Access Theses & Dissertations. 4688.
https://scholarworks.utep.edu/open_etd/4688