Machine-assisted formal reasoning
University of Illinois Urbana-Champaign · Advisor: Prof. Tianyin Xu
I contribute to TLAPS-Bench, a benchmark for evaluating AI systems on completing and constructing machine-checkable TLA+ proofs. My work focuses on the methodology required for credible experiments: scalable proof verification, reliable multi-round evaluation, and accurate measurement of model behavior and cost.
More broadly, I am interested in how machine assistance can make formal reasoning practical for complex software systems without weakening correctness guarantees.
Project repository →
SREGym: AI agents for software reliability
AI agents for Site Reliability Engineering · Current research
I contribute to SREGym, an AI-native platform for developing and evaluating SRE agents in live system environments with realistic cloud-system failures.
My work includes building reproducible Kubernetes incidents, including node conntrack exhaustion that can disrupt an agent's normal diagnostic access path. These scenarios test whether agents can reason about degraded systems and recover them safely.
Project repository →
Bengali speech representations
BUET · Supervisor: Dr. Sadia Sharmin
My undergraduate research examined where Bengali phone-like information emerges across Whisper encoder layers using speaker-disjoint probing, cross-model comparison with XLS-R, and robustness analyses. The resulting paper was accepted to INTERSPEECH 2026.