Text-to-speech is relatively straightforward when the input is a clean, single-language sentence. It becomes considerably harder when the text reflects how people actually speak.
In India, it is possible for one conversation to go from English to a regional language within a single sentence. Numbers can appear as words, digits, IDs, or account details. Names and addresses do not always follow English pronunciation rules. And in a voice conversation, the system has to produce the response quickly enough that the interaction still feels like a conversation.
Indian voice AI aren’t dealing with isolated instances; such cases are included in the input.
Ragini is ConvoZen’s text-to-speech system built around these requirements. It is trained on 800+ hours of proprietary Indian voice data, with an emphasis on conversational and code-switched speech across English, Hindi, Tamil, Telugu, Kannada, and Marathi.
A text-to-speech system does more than map characters to sounds.
For the model to generate speech that is suitable for use in a real conversation, it has to work out how the text should be spoken, for example, where to pause, which words to stress, how numbers should be read, which pronunciation to use for a name or a place, and how the voice should act when the language changes.
Indian everyday speech adds another aspect of complexity: code-switching.
Consider a sentence such as:
“Aapka payment successfully process ho gaya hai, but refund 24 hours mein reflect hoga.”
A system which is designed with monolingual speech in mind might manage to pronounce each word correctly yet still produce an unnatural outcome when sentences contain words from different languages. The issue goes beyond merely covering a range of languages; it is about keeping the pronunciation, prosody, and consistency of the speaker throughout the transition.
It holds structured Indian data as well.
A voice agent might be required to read out a UPI ID, an Aadhaar number, a phone number, an address, a date, or a transaction reference. Since these sequences can never be regarded as ordinary prose, the way they are spoken should take their semantic structure into account.
For voice systems used in production, such details form part of the TTS problem.
Ragini is trained on 800+ hours of proprietary Indian voice acting and conversational data across six languages:
The method used for training is based on spoken conversation rather than studio-style narration. This difference is important since the requirements for a voice agent are different from those for an audiobook or a voiceover process.
Ragini is designed to handle:
The goal goes beyond producing audio that makes sense; it is to narrow the gap between the generated speech and the timing and linguistic patterns that are expected in a real conversation.
In a conversational system, the latency of the text-to-speech system is only one part of the total response time.
A typical pipeline looks like:
User speech → Speech recognition → Language/agent processing → TTS → Audio playback
Ragini is responsible for the final stage, which involves turning the agent’s response into streaming audio.
In the TTS layer, the text is processed with respect to language, code-switching, semantic entities, and speech characteristics before being synthesised. The audio produced can then be streamed to the caller instead of having to wait for the whole response to be generated.
This distinction is important.
A TTS engine which is mainly designed for batch generation can be optimised for audio quality even if it has a higher generation latency. A conversational voice system, on the other hand, has a different requirement: the first piece of audio must be sent quickly enough so that the user feels the interaction to be responsive.
The generation latency and the first-byte performance for Ragini have been designed with this streaming use case in mind.
| Capability | Conventional TTS | Ragini |
| Primary optimization | Narration and general speech synthesis | Real-time conversational speech |
| Training data | Primarily studio-recorded speech | 800+ hours of proprietary Indian voice data |
| Languages | Broad language coverage | English, Hindi, Tamil, Telugu, Kannada, Marathi |
| Code-switching | Often treated as a separate or unsupported case | Designed for mixed-language conversational input |
| Indian formats | May require additional text normalization | Designed to handle formats such as IDs, numbers, and addresses |
| Streaming | Depends on implementation | Streaming generation with first byte under 200ms |
| Deployment | General-purpose applications | Voice AI and real-time conversational systems |
A multilingual text-to-speech system is able to handle six languages on its own. A conversational text-to-speech system must ensure the correct pronunciation, voice characteristics, and intonation when those languages occur together in the same interaction.
The quality of TTS is ultimately determined by the environment in which it functions.
A voice agent needs to work in conjunction with speech recognition, an agent or LLM layer, the telephony infrastructure, and the application logic; even if the quality of the voice is high, it can still result in a bad conversational experience.
Ragini is included within ConvoZen’s wider voice AI system together with Akshara speech recognition and the platform’s agent infrastructure, and ConvoZen is currently handling over 40 million voice AI calls each month.
The benchmark for TTS in that production environment is different from that for offline audio generation since the speech must remain reliable in real conversations, in various languages, under different call conditions, and regardless of response times.
Examples include:
The results cannot be wholly attributed to TTS; they are a result of the overall performance of the voice AI system, and the function of the speech layer is to ensure that the agent’s response is intelligible, delivered in an appropriate manner, and fast enough to maintain the interaction.
The ConvoZen Developer Kit makes Ragini available to teams that are developing their own voice applications.
For real-time conversations, developers can use the streaming APIs, while those needing pre-generated audio should use the batch APIs. The platform offers REST, gRPC and WebSocket interfaces so that the speech layer can be incorporated into existing applications and agent architectures.
Ragini can be used with Akshara STT or instead function as a standalone TTS layer, depending on the application’s architecture.
ConvoZen includes specific resources for Hindi TTS, Tamil TTS, and Kannada TTS for those teams who are developing language-specific experiences.
Measuring the number of languages is an incomplete method of evaluating a TTS system.
For Indian voice applications, a more useful evaluation should consider:
For voice agents, the same thing applies as it does to audio quality in terms of raw quality.
Ragini supports English, Hindi, Tamil, Telugu, Kannada, and Marathi, with a special emphasis on conversational and code-switched speech.
Yes, Ragini has been designed to handle conversational input in multiple languages, including code-switching between Hindi and English within a single utterance.
Ragini has an audio generation latency of under 200 milliseconds, with the first byte being delivered in less than 200 milliseconds.
Yes. Ragini includes support for streaming synthesis and has been designed to function as the TTS layer in real-time conversational systems.
Ragini can be obtained via the ConvoZen Developer Kit and offers streaming and batch APIs through REST, gRPC, and WebSocket. It can be combined with ConvoZen’s other voice AI features or be used in a developer’s own application structure.