Date of Award

2026-05-01

Degree Name

Doctor of Philosophy

Department

Mathematical Sciences

Advisor(s)

Ritwik Bhattacharya

Second Advisor

Tanmay Sen

Abstract

System logs are a primary source of information for identifying faults, misconfigurations, and security incidents in computing infrastructure. As these logs grow in volume and complexity, manual inspection becomes impractical, and automated detection methods are needed. This thesis investigates how compact language models can be applied to the task of log anomaly detection under two operational conditions: when log data is centralized and when it is located at different data sites.

In the first part, we propose LogTinyLLM, which applies parameter-efficient fine-tuning through Low-Rank Adaptation and adapter-based techniques to adapt tiny language models for identifying contextual anomalies in log sequences. We evaluate LogTinyLLM on the Thunderbird and BGL datasets using multiple architectures, including OPT-1.3B, Phi-1.5, DeepSeek-R1-Distill-Qwen-1.5B, and TinyLLaMA-1.1B, across single-module, dual-module, and triple-module LoRA configurations. On Thunderbird, the best performing configuration achieves an F1-score of 0.9857, exceeding existing baselines by up to 32.55 percentage points. Performance remains stable across architectures, with an average F1-score of approximately 0.9845 across all LoRA-based setups.

In the second part, we address the reality that log data in many organizations cannot be pooled together for centralized training. We propose DP-FLogTinyLLM, a federated extension that enables multiple clients to jointly train a shared anomaly detector without exposing their raw logs. The framework combines federated optimization with differential privacy to provide formal privacy guarantees against inference attacks, and uses LoRA to keep the fine-tuning process lightweight enough for resource-constrained client environments. On Thunderbird and BGL, FlogTinyLLM has competitive performances comparable to the centralized settings, with only minor variation in recall on BGL. Also, FLogTinyLLM achieves consistently superior precision and F1 Scores on the Thunderbird dataset compared to the baseline models.

Together, these contributions demonstrate that tiny language models with parameter-efficient adaptation can deliver strong log anomaly detection both in centralized and privacy-preserving federated settings, offering a practical path toward accurate and privacy-aware monitoring of distributed computing systems.

Language

en

Provenance

Received from ProQuest

File Size

126 p.

File Format

application/pdf

Rights Holder

Isaiah Thompson Thompson Ocansey

Share

COinS