04RESEARCH14 models · 3,000+ prompts
Agent Reliability Auditor
Thesis: systematic LLM degradation benchmarking across English, Roman Urdu, and code-switched inputs. 14 models, 3,000+ prompts, 18–34% accuracy drops identified.
— Overview
My final-year thesis: a systematic framework for measuring how LLM output quality degrades outside clean English. The auditor benchmarks models across English, Roman Urdu, and code-switched inputs under adversarial and noisy conditions.
The evaluation harness tests 14 models across 3,000+ prompts in three language modes, and identified statistically significant accuracy drops of 18–34% on code-switched inputs. The write-up has been submitted as a workshop paper and is under review.
— Highlights
- →14 models benchmarked across 3,000+ prompts
- →Three language modes: English, Roman Urdu, and code-switched
- →18–34% accuracy degradation identified on code-switched inputs
- →Submitted as a workshop paper; currently under review