Theses and Dissertations

ORCID

https://orcid.org/0009-0008-7011-5133

Advisor

Zhiqian, Chen

Committee Member

Jingdao, Chen

Committee Member

Gudla, Charan

Date of Degree

5-15-2026

Original embargo terms

Immediate Worldwide Access

Document Type

Graduate Thesis - Open Access

Major

Computer Science (Artificial Intelligence & Robotics)

Degree Name

Master of Science (M.S.)

College

James Worth Bagley College of Engineering

Department

Department of Computer Science and Engineering

Abstract

This thesis extends data contamination auditing for multimodal large language models to multilingual settings. Using LLaVA 1.5 and a high-fidelity French parallel dataset derived from ScienceQA, the study evaluates how performance changes when identical image-question pairs are translated from English to French. The resultsshow a substantial cross-lingual performance decline and frequent flips from correct English predictions to incorrect French predictions, indicating that benchmark performance can depend heavily on memorized English-specific patterns rather than stable multimodal reasoning. To address this weakness, the thesis introduces an inference-time mitigation strategy based on perturbation ensembling and cross-lingual consistency aggregation. The proposed method reduces instance-level leakage without model retraining and offers a practical way to improve the robustness and trustworthiness of multimodal benchmark evaluation. The findings demonstrate the importance of cross-lingual auditing when assessing modern multimodal systems.

Share

COinS