Agharia Kasimali
Multilingual ASR Research (Annam.ai & IIT Ropar Initiative)
This space acts as a centralized verification archive for the automatic speech recognition (ASR) modules developed during my research internship at Annam.ai (Government of India & IIT Ropar Initiative) from August to October 2025.
The work focuses on multilingual ASR for 10+ Indian languages, analyzing language-specific acoustic characteristics (Indo-Aryan, Dravidian families) and speaker variability across regions and demographics, along with fairness and scalability considerations.
🎥 ASR Models & Artifacts
Implementation Note: As these modules were originally developed and executed in Google Colab, the interactive Gradio web application is currently only deployed and actively running under the Whole Pipeline link. The remaining links serve as static verification archives containing the raw code, evaluation scripts, and fine-tuning logs.
Click on any link below to review model demos, fine-tuning experiments, and pipeline components directly in your browser:
1. Multilingual ASR Demos
2. Fine-tuning & Experiment Logs
- FINETUNINNG — ASR fine-tuning experiments and training logs.
- FINETUNNING WHISPER — Whisper-based fine-tuning runs for Indian languages.
- XLS-R1B — XLS-R 1B parameter model experiments for multilingual ASR.
3. Pipeline Components & Utilities
-
Whole Pipeline
● Active Gradio App
— End-to-end ASR inference pipeline (audio → text).
- INDICTRANS2 — IndicTrans2 integration for ASR output translation.
- IndicXlitF — Romanized-to-native script conversion for ASR outputs.
🔒 Note: This archive is intended for research verification and demonstration purposes. Model weights, datasets, and internal configurations follow the respective institutional and licensing guidelines.