Date of Award

2026-05-01

Degree Name

Master of Science

Department

Mathematical Sciences

Advisor(s)

Nilotpal Sanyal

Abstract

Modern cancer research increasingly relies on machine learning methods to model complex relationships between high-dimensional genomic predictors and clinically relevant outcomes. Gradient boosting methods, particularly XGBoost, have become widely used due to their strong predictive performance and ability to capture nonlinear effects and interactions. However, these methods are sensitive to contamination in the data, especially in the response variable, where even a small fraction of aberrant observations can disproportionately influence model fitting and degrade generalization.

In this thesis, we propose a contamination-aware extension of XGBoost, termed CA-XGBoost, which improves robustness by explicitly modeling residuals through a latent mixture framework. This formulation distinguishes between clean and contaminated observations and yields adaptive, observation-specific weights that are iteratively updated during training. These weights are incorporated directly into the boosting procedure, effectively downweighting high-residual observations while preserving the underlying signal. The approach retains the flexibility and computational efficiency of standard XGBoost while enhancing its robustness.

The proposed method is evaluated through extensive simulation studies and a real breast cancer gene expression application using TCGA PanCancer Atlas data. Simulation results show that CA-XGBoost consistently achieves improved predictive accuracy and stability under contaminated training conditions. In the real data analysis, using MKI67 expression as a continuous response, the method demonstrates superior performance compared to standard XGBoost when contamination is present, yielding lower prediction error and more stable residual behavior. These findings highlight the practical value of incorporating contamination awareness into boosting frameworks for reliable predictive modeling in complex biomedical settings.

Language

en

Provenance

Received from ProQuest

File Size

66 p.

File Format

application/pdf

Rights Holder

Francis Opoku

Share

COinS