India has hundreds of languages and a metric ton of dialects to go with them. As you travel across the country, you notice subtle (and sometimes not-so-subtle) shifts in how a language is spoken. Dialect information improves ASR (Automatic Speech Recognition / Text To Speech) performance especially for low resource languages, which most Indian languages are.[1][2][3]

This project aims to pinpoint a speaker’s dialect, using their district of origin as a proxy.

However, two things to keep in mind:

  • People move! Inter-district migration is pretty common, which means a speaker’s current location might not reflect their native dialect.
  • There are no prior studies on dialect identification with districts as a proxy, thus we do not have a baseline to compare to.

Vaani: A Massive & Diverse Dataset of Indian Speech

Our experiments are done on the Vaani dataset, an extensive collection of speech samples gathered across India.

Vaani map with colored states
The colors on the map indicate the recorded duration in hours for each state.

These stats are severely outdated. Here are the updated stats.

Some quick stats on the dataset to get a feel for it:

  • Massive Scale: Roughly 16,000 hours of speech data and increasing rapidly.
  • Rich Diversity: Covers 54 Indic languages from 80 districts across 12 states.
  • Valuable Metadata: Includes district of recording, speaker demographics (where available), and language labels.
  • Transcriptions: A small portion ~5% comes with manual transcriptions.
  • The Team Behind It: Developed by ARTPARK, IISc Bangalore, and Prof. Prasanta Kumar Ghosh (EE Dept, IISc), with funding from Google.

For our experiments, we carve out specific subsets from Vaani, focusing on particular states or languages, and use the district-level labels to train our models. We use a 80-10-10 split stratified by speakers balanced by district (dialect) counts.

Our Approach: Finetuning Various Speech Models

We finetuned several pretrained models on this task and compared their performance on various subsets.

Wav2Vec2 and its Variants by Meta AI

Wav2Vec2 is a popular framework that learns speech representations from raw audio in a self-supervised way. We experimented with:

  • Wav2Vec2-Base / Large: Pre-trained mainly on English speech.
  • Wav2Vec2-XLS-R (300M, 1B, 2B): Pre-trained on diverse multilingual speech data. This is a key difference!
  • MMS (300M, 1B): Meta AI’s Massively Multilingual Speech models, trained on thousands of languages.
  • Wav2Vec2-Conformer: Incorporates Conformer blocks for potentially better architectural performance.

We finetuned them for dialect classification by adding a classification head at the last layer.

WavLM by Microsoft

WavLM is designed to understand both spoken content (for ASR) and paralinguistic tasks (speaker identity, emotion recognition and the sort). This makes it a strong candidate for dialect identification, which relies on subtle acoustic and phonetic cues. We primarily used the WavLM-Base model. (Experiments with WavLM were conducted by Bingimalla Yasaswini from Vision and AI Lab at IISc as part of this project.)

Vakyansh by EkStep Foundation

Vakyansh offers Wav2Vec2 models pre-trained on English and then fine-tuned for specific Indian languages. We used a model fine-tuned on 96 hours of Telugu data, which itself was derived from their Hindi fine-tuned model. The base architecture is Wav2Vec2-Base. (Experiments with Vakyansh were conducted by Bingimalla Yasaswini.)

Experiments and What We Found

We ran several experiments, focusing on different slices of the Vaani dataset.

Zooming in on Andhra Pradesh / Telugu Districts with Wav2Vec2

We picked 6 Telugu-speaking districts in Andhra Pradesh: Anantpur, Chittoor, Guntur, Krishna, Srikakulam, and Vishakapattanam.

  • Wav2Vec2-Base and Large (pre-trained on English) performed no better than random chance (around 17% ~ 100/6% accuracy). This highlighted that English pre-training wasn’t cutting it for the nuances of Telugu dialects.
  • Wav2Vec2-XLS-R-300m pre-trained on diverse languages, achieved a much better validation accuracy of 53% after just 5 epochs of fine-tuning! This strongly suggested that multilingual pre-training is crucial for this task.
  • MMS-300m performed similarly to XLS-R, also hitting 53% validation accuracy.
  • Wav2Vec2-Conformer didn’t show significant improvement here, reaching only 35% accuracy. This might be due to its English pre-training and the need for more extensive multilingual fine-tuning. It was also slower to train due to its large size.

Here’s a look at the confusion matrix for the XLS-R model on the 6 Andhra Pradesh districts, alongside a map showing their locations. Interestingly, we didn’t find a clear link between geographically proximity of the districts are and model confusion.

Confusion Matrix for Wav2Vec2-XLS-R on 6 AP districts
Confusion matrix for Wav2Vec2-XLS-R-300m on 6 Andhra Pradesh districts (Telugu).
Map of Andhra Pradesh districts
Map highlighting the Andhra Pradesh districts used in the experiment (teal).

Andhra Pradesh / Telugu Districts with WavLM

Fine-tuning the WavLM-Base model for 20 epochs yielded a validation accuracy of 39% for the 6 Andhra Pradesh Telugu districts.

The confusion matrix shows that, again, no obvious correlation with geographic proximity was observed.

Confusion Matrix for WavLM on 6 AP districts
Confusion matrix for WavLM on 6 Andhra Pradesh districts (Telugu) - Validation Set.
Map of Andhra Pradesh districts
Map highlighting the Andhra Pradesh districts used in the experiment (teal).

Vakyansh Telugu Model Results

Using the Vakyansh Wav2Vec2 model (fine-tuned for Telugu), we again targeted the 6 Andhra Pradesh districts, achieving a validation accuracy of 29%.

Confusion Matrix for Vakyansh on 6 AP districts
Confusion matrix for the Vakyansh Telugu model on 6 Andhra Pradesh districts.
Map of Andhra Pradesh districts
Map highlighting the Andhra Pradesh districts used in the experiment (teal).

Tackling Hindi Language Districts

In our next experiment, we broadened our scope to 46 districts across various North Indian states where Hindi is prevalent. Using Wav2Vec2-XLS-R, we achieved an accuracy of about 35% on the validation set. This task was tougher, likely due to the larger number of classes and the linguistic diversity even within Hindi-speaking regions.

The districts included were: Araria, Begusarai, Bhagalpur, Darbhanga, East Champaran, Gaya, Gopalganj, Jahanabad, Jamui, Kishanganj, Lakhisarai, Madhepura, Muzaffarpur, Purnia, Saharsa, Samastipur, Saran, Sitamarhi, Supaul, Vaishali, Budaun, Deoria, Etah, Ghazipur, Gorakhpur, Hamirpur, Jalaun, Jyotiba Phule Nagar, Muzaffarnagar, Varanasi, Churu, Nagaur, Tehri Garhwal, Uttarkashi, Balrampur, Bastar, Bilaspur, Jashpur, Kabirdham, Korba, Raigarh, Rajnandgaon, Sarguja, Sukma, Jamtara, and Sahebganj.

A Diverse Multi-Language Challenge with WavLM

To test how well WavLM generalizes across more distinct linguistic and geographic lines, we picked 12 districts, each from a different state (and likely a different language). The labels were: Anantpur, Araria, Balrampur, NorthSouthGoa, Jamtara, Belgaum, Karimnagar, Aurangabad, Churu, Budaun, TehriGarhwal, and DakshinDinajpur.

Performance varied quite a bit, with some districts like Belgaum being identified more accurately.

Confusion Matrix for WavLM on 11 diverse districts
Confusion matrix for WavLM on a diverse set of 11 districts.

WavLM on Bihar Districts

We also tested WavLM on 20 districts from Bihar (primarily Hindi-speaking). After filtering the dataset to include female speakers with over 20 years of residency in their respective districts, the model achieved an accuracy of 26%.

Confusion Matrix for WavLM on Bihar districts
Confusion matrix for WavLM on Bihar districts (female speakers, >20 years residency).

Takeaways

  • Multilingual Pre-training is Key: Models pre-trained on diverse languages (XLS-R, MMS, WavLM) consistently outperformed those trained primarily on English.
  • Moderate Success: We achieved accuracies in the 30-50% range for 6-class problems, which is promising but shows there’s room for improvement.
  • Data Filtering Helps: Filtering audio samples to be at least 5 seconds long boosted performance by about 10% for Wav2Vec2-XLS-R on the Telugu 6-district classification task.
  • No Simple Geographic Link: For the Andhra Pradesh case, we didn’t see a strong connection between geographic proximity and confusion. This could mean dialect boundaries are more complex than just adjacency, or that models might be picking up on other speaker or recording-specific details rather than pure dialectal features.

Future Directions

If you were to continue this line of research, I’d suggest the following:

  • Different kind of models: Try models with a different pre-training objective/architectures like Whisper (known for strong multilingual abilities).
  • Indic Language Specialists: Using models specifically pre-trained on large Indic language corpora, like IndicWav2Vec.
  • ASR as an Auxiliary Task: Training dialect identification as a secondary goal to ASR. This involves prepending transcripts (<telugu_chittoor> transcript...) and training an end-to-end ASR model. The idea is that forcing the model to understand the content might help it learn dialectal features more effectively.

Conclusion

Our journey into district-level dialect identification in India using the Vaani dataset has been revealing. Multilingual pre-trained models like Wav2Vec2-XLS-R and MMS show the most promise so far, achieving respectable accuracies (around 50-53%) for a 6-district Telugu classification task. However, the challenge is significant, with considerable confusion between dialects. Future efforts will focus on leveraging cutting-edge multilingual models, exploring joint ASR and dialect ID training, and potentially using multi-stage approaches to better isolate those subtle dialectal cues in speech.


Acknowledgements: We thank Prof. Venkatesh Babu for his guidance and support, and VAI Lab for providing the compute resources. I thanks Bingimalla Yasaswini for her work on the WaveLM and Vakyansh model experiments and her valuable feedback. We also thank the creators of the Vaani dataset (ARTPARK, IISc, Google) and the developers of the open-source models and libraries used. Special thanks to Saurabh Kumar from SPIRE lab for his insightful inputs.

References

  1. Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion
  2. Improving Automatic Speech Recognition with Dialect-Specific Language Models
  3. Building Robust and Scalable Multilingual ASR for Indian Languages

Notes in post (Jan 12th 2026): This was my first foot into deeplearning, the project was tough, I had scarce resources, sharing and fighting over a Turing GPU and a 9th gen consumer cpu chip which took 5 minutes just to load torch. I know the numbers aren’t impressive but this is the first work of its kind and there are a number of things I’d try to improve the performance, starting with a deeper error analysis and saliency analysis on the waveform/spectrogram.