Theses and Dissertations

ORCID

https://orcid.org/0009-0006-5099-0091

Advisor

Zhiqian, Chen

Committee Member

Jingdao, Chen

Committee Member

Gudla, Charan

Date of Degree

5-15-2026

Original embargo terms

Immediate Worldwide Access

Document Type

Graduate Thesis - Open Access

Major

Computer Science (Artificial Intelligence & Robotics)

Degree Name

Master of Science (M.S.)

College

James Worth Bagley College of Engineering

Department

Department of Computer Science and Engineering

Abstract

This thesis evaluates bias and harmful language generation in five open-source language models and tests practical mitigation methods that do not require retraining. Two masked models are assessed with a sentence-pair benchmark for stereotype preference, and three generative models are assessed with a prompt-based benchmark for harmful continuations across demographic domains. The study uses a unified experimental workflow to compare model behavior, summarize differences across bias categories, and measure changes after intervention. Results show that the masked models favor stereotypical content above a random baseline, while the generative models usually produce low average toxicity but still show uneven risk across domains. Two post-hoc mitigation methods, score-based filtering and safety-oriented prompting, reduce harmful outputs while largely preserving usable text generation. The findings support lightweight mitigation as a practical strategy for improving safer deployment of open-source language models.

Share

COinS