Date of Award

2026-05-01

Degree Name

Master of Science

Department

Computer Science

Advisor(s)

Anantaa Kotal

Abstract

Machine learning systems deployed in high-stakes domains are increasingly expected to satisfy demands beyond predictive accuracy-including fairness across demographic groups, protection of sensitive information, and explanations that human stakeholders can inspect and trust. This thesis investigates how those demands can be met through learning frameworks that explicitly govern the relationship between data and models, arguing that trustworthiness is a design problem rather than a post hoc correction. The thesis is organized around three studies, each targeting a distinct point of data-facing control. The first develops CondFairGen, a fairness-aware conditional generator for tabular data that improves subgroup equity by dynamically reweighting the group conditions that receive exposure during training, rather than modifying the generator architecture or loss function. Evaluated across five benchmark datasets and five downstream classifiers, CondFairGen with intersectional exposure control achieves the best combined fairness-?utility trade-off among compared methods. The second introduces FAIRPLAI, a human-in-the-loop framework that makes the fairness-privacy-accuracy paradox explicit and navigable: it constructs privacy-fairness frontiers showing achievable trade-offs, formalizes stakeholder requirements through a policy tuple, and provides a bidirectional translation layer connecting plain-language goals to enforceable technical constraints. Evaluated on five benchmark datasets under differential privacy, FAIRPLAI satisfies joint fairness and accuracy requirements in three of five settings and achieves 86% upward and 91% downward translation fidelity. The third develops a grammar-guided symbolic regression system using Monte Carlo tree search and bounded LLM-suggested production rules, treating the grammar itself as a governed representation that determines which interpretable structures can be discovered from data. Preliminary experiments on Arrhenius and Kepler benchmarks demonstrate the architecture and the auditability of accepted and rejected rules. Taken together, these projects support a unified claim: trustworthy machine learning requires frameworks that explicitly regulate the representation, exposure, and use of data. Fairness, privacy, and interpretability are not independent failure modes to be patched after training - they are consequences of design choices made throughout the learning process.

Language

en

Provenance

Received from ProQuest

File Size

73 p.

File Format

application/pdf

Rights Holder

David Anthony Sanchez

Share

COinS