HA
All projects
04RESEARCH14 models · 3,000+ prompts

Agent Reliability Auditor

Thesis: systematic LLM degradation benchmarking across English, Roman Urdu, and code-switched inputs. 14 models, 3,000+ prompts, 18–34% accuracy drops identified.

— Overview

My final-year thesis: a systematic framework for measuring how LLM output quality degrades outside clean English. The auditor benchmarks models across English, Roman Urdu, and code-switched inputs under adversarial and noisy conditions.

The evaluation harness tests 14 models across 3,000+ prompts in three language modes, and identified statistically significant accuracy drops of 18–34% on code-switched inputs. The write-up has been submitted as a workshop paper and is under review.

— Highlights

  • 14 models benchmarked across 3,000+ prompts
  • Three language modes: English, Roman Urdu, and code-switched
  • 18–34% accuracy degradation identified on code-switched inputs
  • Submitted as a workshop paper; currently under review