Date of Award

2026-05-01

Degree Name

Master of Science

Department

Computational Science

Advisor(s)

Jonathon E. Mohl

Abstract

Prediabetes is a critical health condition that increases the risk of developing type 2 diabetes. The hemoglobin A1c (HbA1c) test diagnoses patients with prediabetes, but the disease has already caused metabolic alterations. Early detection is essential for timely interventions, and machine learning models offer a promising approach to identify prediabetic individuals through the analysis of biomarkers such as cytokines. We compared four classifiers-logistic regression, decision tree, random forest, and k-nearest neighbors - using cytokines (TNF-α, MCP-1, IL-1β, IL-6, IFN-γ), age, BMI, and waist-to-hip ratio (WHR). Models were evaluated using stratified 5-fold cross-validation and ROC-AUC. K-Nearest Neighbors (k-NN) achieved the highest performance (CV AUC = 0.771 ± 0.037), closely followed by random forest (0.750 ± 0.026), logistic regression (0.704 ± 0.027), and decision tree (0.637 ± 0.051). Feature importance varied by model: decision tree and random forest highlighted age, MCP-1, TNF-α, IL-6, and WHR as key predictors. IL-1β, IFN-γ, and BMI showed negligible importance. Random forest provided the best balance of predictive performance, sensitivity, and interpretability. These findings support the use of non-linear machine learning models for integrating clinical and biomarker data in chronic disease risk assessment.

Language

en

Provenance

Received from ProQuest

File Size

46 p.

File Format

application/pdf

Rights Holder

Luisa Veronica Gracia Mazuca

Share

COinS