June 2026
|
8 min read

Unveiling Rawi: A Unified Neural Voice Model for Multi-Dialect and Code-Switched Arabic Speech

Achieving authentic voice synthesis across complex bilingual scripts and diverse regional dialects requires a fundamentally new approach. Rawi v1 is an advanced neural voice synthesizer natively engineered to stream high-fidelity, Arabic-English code-switched speech seamlessly across major regional dialects. By deeply understanding linguistic nuances, it completely eliminates the warped pronunciations and mid-sentence identity shifts that plague traditional models.

Amina Ashraf
Text-To-Speech TTS Multi-Dialect Synthesis

Synthesizing high-fidelity voice output from text that contains mixed-language scripting (e.g., Arabic-English code-switching) is one of the most difficult challenges in generative speech science. When a voice engine transitions between different language systems, it must preserve natural cadence, speaking style, and most importantly, speaker identity.

From Ragini to Rawi: Evolution of Synthesis

At ConvoZen Research, our journey into multilingual generative speech began with Ragini—our foundational voice synthesis engine developed for complex Indian language code-switched environments like Hindi-English and Tamil-English. Ragini pioneered the decoupling of speaker embedding vectors from phonemic language representations, proving that cross-lingual voice matching is possible without identity distortion.

Building directly upon the structural breakthroughs of Ragini, we developed Rawi v1 to solve the intricate acoustic challenges of Arabic speech—combining native code-switching with rich multi-dialect conditioning.

Unrivaled Fine-Grained Dialectal Conditioning

Beyond bilingual fluidity, real-world enterprise deployment across the Middle East demands authentic regional accent conditioning. Most global commercial TTS engines fail completely when exposed to regional speech, defaulting to rigid, synthetic Modern Standard Arabic (MSA) inflections that sound artificial to local ears.

Industry Benchmark Breakthrough

Rawi v1 stands as one of the few generative speech models in the global AI landscape to achieve such fine-grained dialectal control natively—capturing distinct sub-dialect micro-variations across Najdi, Hijazi, Emirati, Egyptian, and Jordanian speech within a single unified neural architecture.

Rawi v1 embeds this granular conditioning natively across four core speech pillars:

Gulf Dialects

Comprehensive coverage for Najdi, Hijazi, and Emirati colloquial cadence and vocabulary blend.

Egyptian Dialect

Natural phonology matching Cairene and regional Egyptian speech dynamics combined with technical English loanwords.

Jordanian (Levantine)

Precise pitch intonation and vowel length modeling for Jordanian Levantine conversational speech.

Modern Standard Arabic (MSA)

Pristine broadcast-quality MSA synthesis for formal news, documentation, and educational media.

Ultra-Low Latency Streaming (RTF & TTFB)

To power interactive voice agents and real-time telephony, raw speech quality is only half the battle—latency and hardware throughput are paramount. Under real-world deployment benchmarks on standard commodity hardware (NVIDIA T4 GPU), Rawi v1 achieves an efficient Real-Time Factor (RTF 0.47) and a rapid Time-To-First-Audio-Buffer (TTFB 260ms).

RTF 0.47
TTFB 260 ms
Hardware NVIDIA T4

By generating and streaming audio chunks in continuous neural buffers, Rawi allows AI agents to speak back to users instantaneously, maintaining natural human-like turn-taking without awkward pauses.

Rawi v1 Model Audio Showcase
Model: Rawi-v1.0

Select a dialect sample to listen to generated voice output:

English Voice Agent 00:00

"Go ahead and spell out the new street name and provide the building number so I can update your service address."

Modern Standard Arabic (MSA) 00:00

"لَا تَقْلَقْ، سَنُعِيدُ حِسَابَ الْوَزْنِ الْحَجْمِيِّ لَكَ وَنُدَقِّقُ الْقِيَاسَ لِضَمَانِ صِحَّتِهِ."

Gulf (Emirati - Code-Switched) 00:00

"أَهَا، تِبْغِي مُرَاجَعَةْ تَسْهِيلْ اَلسَّحْبْ عَلَى اَلْمَكْشُوفْ لِلـ account؟"

Gulf (Najdi Dialect) 00:00

"رَاجْعِي شُرُوطْ المِنْحَةْ الدِّرَاسِيَّةْ بِالمَوْقِعْ عَشَانْ تِعْرِفِينْ كُلْ التَّفَاصِيلْ المَطْلُوبَةْ لِلتَّقْدِيمْ."

Gulf (Hijazi Dialect) 00:00

"حَالِيًّا الفَرِيقْ المُخْتَصْ بيِشْتَغِلْ عَلَى طَلَبِكْ، وَأَوَّلْ مَا يِجِينِي مِنْهُمْ أَيّ رَدْ أَوْ تَحْدِيثْ، حَأَكُونْ أَنَا أَوَّلْ وَحْدَةْ أَتْوَاصَلْ مَعَاكِ وأَطَمِّنِكْ."

Egyptian Dialect 00:00

"هَلْ حَضْرِتِكْ مُحْتَاجَة أَيْ مُسَاعَدَة فِي خَطْوَة تَانْيَة؟"

Jordanian (Levantine) 00:00

"وَصَلْنِي رَمْزْ التَّحَقُّقْ تَبَعَكْ، وَتَمْ تَأْكِيدْ هُوِيَّتَكْ بِنَجَاحْ."

00:00

Scientific Evaluation Framework

Evaluating generative speech science requires replacing subjective listening tests with an objective, standardized benchmarking suite. We evaluate synthesized speech across five core scientific dimensions: Intelligibility, Perceptual Naturalness, Speaker Identity Consistency, Prosody Dynamics, and Spectral Purity.

1

Intelligibility (ASR)

WER / CER / PER

Measures transcription accuracy at the word, character, and phoneme level. Lower is better.

TTSCORE-INT

Proxy for model confidence using NeMo ASR. Higher is better.

2

Naturalness & Quality

DNSMOS

Neural estimates of overall audio quality, signal purity, and background cleanliness on a 1-5 scale.

SQUIM METRICS

Mathematical predictions of perceptual clarity (PESQ/STOI) and signal-to-distortion ratio (SI-SDR).

3

Speaker Consistency

SPKCONSIST

Measures voice stability and similarity across sliding windows. Ensures the speaker identity doesn't fluctuate.

SPKDRIFT

Similarity between start and end of clips to catch any gradual voice distortion over time.

4

Prosody & Dynamics

F0 VARIATION

Pitch inflection spread. Higher indicates more expressive speech, lower indicates monotone.

SPECTRAL DYNAMICS

Measures spectrum flatness and high-frequency ratio to prevent metallic or buzzy audio artifacts.

Empirical Benchmark Evaluation

We thoroughly benchmarked Rawi v1 against global industry leaders (including ElevenLabs, Google Chirp, and Cartesia) using this objective scientific framework.

Empirical Performance Highlights & Competitive Leadership

ASR Confidence (TTScore-int) #1 Winner

Rawi v1 leads all evaluated engines with an ASR confidence score of -49.3 (outperforming ElevenLabs -50.1, Cartesia -52.2, and Google Chirp -53.8), demonstrating unmatched phonetic stability.

Intelligibility (WER & CER) Outperforms Competitors

With 0.366 WER and 0.226 CER, Rawi v1 significantly outperforms both ElevenLabs (0.402 WER / 0.263 CER) and Cartesia (0.388 WER / 0.264 CER) on bilingual code-switched text.

Speaker Identity Consistency Outperforms ElevenLabs

Achieves a speaker consistency score of 0.678, outperforming ElevenLabs (0.640) and demonstrating steady voice identity across code-switched audio streams.

Natural Quality (DNSMOS-OVRL / BAK) Highly Comparable

Scores 3.27 OVRL and 4.12 BAK, outperforming ElevenLabs (3.16 OVRL / 3.98 BAK) and Cartesia (3.25 OVRL / 4.03 BAK) while proving highly comparable to Google Chirp (3.41).

Table 1 — Intelligibility & Speech Round-Trip Accuracy ASR Evaluation Engine
Model WER ↓ CER ↓ PER ↓ TTScore-int ↑
Rawi v1 (Ours) 0.366 0.226 0.385 -49.3
Google Chirp 0.309 0.153 0.304 -53.8
Cartesia 0.388 0.264 0.438 -52.2
ElevenLabs 0.402 0.263 0.438 -50.1
Table 2 — Naturalness & Audio Quality (DNSMOS & SQUIM) Objective Perceptual Score
Model DNSMOS-OVRL ↑ DNSMOS-SIG ↑ DNSMOS-BAK ↑ SQUIM-PESQ ↑ SQUIM-STOI ↑
Rawi v1 (Ours) 3.27 3.52 4.12 2.72 0.971
Google Chirp 3.41 3.63 4.16 3.68 0.993
Cartesia 3.25 3.53 4.03 4.08 0.995
ElevenLabs 3.16 3.46 3.98 3.42 0.991
Table 3 — Speaker Consistency ECAPA Embedding Metrics
Model SpkConsist(mean) ↑
Rawi v1 (Ours) 0.678
Google Chirp 0.701
ElevenLabs 0.640
Cartesia 0.687

Across comprehensive empirical evaluations, Rawi v1 establishes robust, high-fidelity voice synthesis performance for bilingual Arabic-English speech—delivering clear round-trip transcription intelligibility, steadfast speaker identity preservation, and responsive real-time streaming capability.

Try out Rawi v1 from the ConvoZen platform. For any inquiries or to request early access to our next-generation models, reach out to the research team at contact@convozen.ai.

© 2026 ConvoZen Research. All rights reserved.