Gujarati Speech-to-Text: Accuracy Benchmark for Telephonic Speech

See how ConvoZen’s Akshara performs on Gujarati speech-to-text, including real-world telephonic audio, WER, CER, and benchmark comparisons.
Book Demo
Gujarati Speech-to-Text Benchmark ResultsWhy Telephonic Accuracy Matters More Than Open-Domain AccuracyWhat This Means for Your Contact CentreConvozen’s Broader Speech AI StackFAQs

Contact centres serving Gujarat, and Gujarati-speaking customers across Mumbai, Surat, and the diaspora markets in the Middle East, run into the same wall: most speech recognition models are tuned for English and Hindi first, and Gujarati accuracy drops sharply the moment a call moves from a clean studio recording to a real customer on a noisy mobile line. Misheard policy numbers, garbled account details, and broken transcripts in QA dashboards are the direct result.

Convozen’s speech-to-text model, Akshara, was benchmarked on Gujarati speech-to-text as part of a wider evaluation covering 9 Indian languages, run against Sarvam Saaras v3 and ElevenLabs Scribe v2 across 16 hours of audio, including real-world telephonic conditions.


Gujarati Speech-to-Text Benchmark Results

The table below reports Word Error Rate (WER) and Character Error Rate (CER) on Gujarati audio, combining the public Indic Voices + Vaani dataset and Convozen’s own Indic Telephonic Voice Bench. Lower is better for both metrics.

Metric Akshara Sarvam Saaras v3 ElevenLabs Scribe v2
WER (all datasets) 17.09% 22.83% 32.75%
CER (all datasets) 11.21% 17.08% 23.58%
WER (Indic Voices + Vaani, open domain) 12.91% 16.52% 36.32%
WER (telephonic speech) 31.41% 45.54% 78.89%

Source: Akshara ASR Benchmark Report, Convozen, February 2026. 0.63 hours of Gujarati evaluation audio on the combined benchmark; methodology and full per-language results available in the report.

On the combined benchmark, Akshara’s Gujarati WER is 25% lower than Sarvam Saaras v3 and 48% lower than ElevenLabs Scribe v2. On telephonic audio specifically, the gap widens further: Akshara holds a 31.41% WER against ElevenLabs Scribe v2’s 78.89%, a benchmark condition designed to reflect real contact centre calls rather than studio-clean speech.


Why Telephonic Accuracy Matters More Than Open-Domain Accuracy

Most public ASR benchmarks are built on read speech or open-domain recordings. Real contact centre audio looks nothing like that: 8kHz telephony compression, background noise, agents and customers talking over each other, and code-switching between Gujarati and English mid-sentence. Convozen built the Indic Telephonic Voice Bench specifically to test models under these conditions, using 7.24 hours of real-world telephonic audio across the same 9 languages.

The gap between Akshara’s open-domain and telephonic Gujarati WER (12.91% to 31.41%) reflects how much harder telephonic speech is across every model tested, and every competitor’s error rate widens by a larger margin than Akshara’s on the same shift. This is the number that matters for a contact centre evaluating transcription accuracy on live customer calls, not open-domain benchmark audio.


What This Means for Your Contact Centre

Gujarati is currently part of Akshara’s evaluated language set rather than a separately packaged, generally available API in the way Convozen’s five core languages, English, Hindi, Tamil, Telugu, and Kannada, are. The benchmark results above reflect model-level performance, tested under the same conditions as Convozen’s confirmed languages. Convozen also runs a separate benchmarked track for Marathi speech-to-text, evaluated under the same methodology.

If your contact centre handles Gujarati-speaking customers at volume, whether in BFSI, insurance, or D2C support, this is the accuracy data to work from when evaluating a speech AI vendor. Teams already running Convozen for Hindi or English transcription can raise Gujarati coverage directly with their account team to discuss current availability and rollout timelines.


Convozen’s Broader Speech AI Stack

Akshara is one half of Convozen’s speech layer. On the output side, Ragini, Convozen’s text-to-speech engine, generates natural-sounding voice responses in under 200ms, trained on over 800 hours of proprietary Indian voice acting. Both engines sit inside Convozen’s MSOC (Multi-Session Omni-Channel) architecture, which keeps conversational context persistent across calls, chat, and WhatsApp. So, a Gujarati-speaking customer’s history carries across channels instead of resetting with every interaction.

Convozen processes 40M+ voice AI calls and audits 50M+ conversations monthly, giving the underlying models continuous exposure to real telephonic speech, including cross-lingual and code-switched audio, rather than static benchmark sets alone.

Reach out to our team today to discuss Gujarati language coverage and rollout timelines for your contact centre.


FAQs

1. How was the Gujarati WER calculated?

WER and CER were computed at the utterance level across two benchmarks, an open-domain public dataset and Convozen’s own telephonic speech bench, then aggregated per language.

2. Does Akshara handle Gujarati-English code-switching?

The Indic Telephonic Voice Bench includes conversational, code-switched audio by design, since that pattern is common in real Indian contact centre calls.

Didn’t find what you’re looking for?Write to us at contact@convozen.ai
Ready to decode AI‑powered conversations?Get Started
Ready To Deploy Your Agentic Workforce?See ConvoZen In Action In Your Environment
Schedule Demo