The new benchmark combines speech recognition and speaker identification evaluation using 108 hours of real-world conversations, covering multiple speakers, dialects, acoustic environments and code-mixed speech across India.
Sarvam AI, in partnership with AI4Bharat, has launched Indic DiarBench, a benchmark dataset aimed at evaluating automatic speech recognition (ASR) and speaker diarization together across all 22 scheduled Indian languages. The dataset contains approximately 108 hours of naturally occurring, multi-speaker conversations and has been made available through Hugging Face.
The initiative seeks to address a limitation in conventional speech AI evaluation, where transcription and speaker identification are generally assessed as separate capabilities. ASR converts spoken language into text, while diarization identifies which person spoke at a particular point. In practical conversations, however, these functions need to operate together, especially when speakers interrupt one another, respond simultaneously or use conversational backchannels.
Designed for real-world Indian speech
Indic DiarBench incorporates conversations involving between two and nine speakers, with recordings featuring interruptions, rapid exchanges, overlapping speech and varied conversational patterns. The benchmark combines meeting recordings captured in controlled environments with more spontaneous conversations sourced from YouTube, providing a broader representation of real-world audio conditions.
Around 81 hours of the dataset comprise meeting conversations, while nearly 28 hours come from in-the-wild recordings. The meeting dataset includes 485 individual speakers from 189 districts across urban and rural areas of India, bringing variations in dialects, educational backgrounds and speaking styles into the evaluation set.
The meeting recordings include approximately 53 hours of near-field audio and 27 hours of far-field recordings captured using distant, conference-style microphones. The latter introduces challenges such as ambient noise and reverberation. The in-the-wild segment focuses on the 10 most widely spoken Indian languages, with approximately two hours of data per language and around 750 speakers.
Multi-stage annotation and evaluation
According to Sarvam AI, the dataset underwent a five-stage annotation workflow. Multiple ASR systems were initially used to generate preliminary transcripts, which were subsequently reviewed by human annotators for word accuracy, timestamps and speaker attribution. Further quality checks examined speaker labels, code-mixed speech, numerals, timestamps and non-speech events.
The benchmark also reflects India's multilingual communication patterns by including conversations where English is mixed with Indic languages. Transcriptions are provided in native Indic scripts, while English words are represented in Roman script and numbers in Arabic numerals. An expert review was conducted as the final quality-control stage, with files requiring corrections sent back for revision.
Indic DiarBench enables assessment through three key metrics: Diarization Error Rate (DER), concatenated minimum-permutation Word Error Rate (cpWER) and Word Diarization Error Rate (WDER). These respectively measure speaker identification accuracy, combined transcription and speaker attribution performance, and the accuracy of assigning individual words to the correct speaker.
The complete dataset, annotations and evaluation methodology are available on Hugging Face. Sarvam AI and AI4Bharat are also scheduled to present the benchmark at Interspeech 2026, alongside a research paper published on arXiv.
