Streaming, Contextual and Conversational Speech Models for Africa
Our earlier research agenda explored streaming transcription, efficient audio representations, adaptive speech generation and duplex conversation. Recent work makes several of these directions more concrete. But faster speech models do not automatically understand African languages, and expressive voices do not automatically produce useful conversations.
Consider a customer saying, “I don transfer the money, but e never reflect.” The model needs to preserve the Pidgin, understand the payment complaint, and distinguish it from confirmation that a payment succeeded. Add another person speaking nearby, a local name and a correction to an amount, and the problem becomes much larger than transcribing a clean sentence.
Realtime Multilingual Transcription
Why does streaming remain difficult? A model has to produce text before it has heard the full utterance. More future audio can resolve ambiguity, but waiting for it adds delay. Emitting text earlier can introduce revisions, especially around short words, names and language switches.
Voxtral Realtime addresses this through a causal audio encoder and aligned audio-text streams during training. Qwen3-ASR offers smaller recognition models alongside a larger variant. These are useful advances beyond treating every incoming chunk as another offline transcription.
However, a fast first token does not mean the transcript is already reliable. A caller may correct a delivery location halfway through a sentence or switch into Yorùbá after giving an amount in English. The important tradeoff is between delay and stable meaning, not just delay and the first visible word.
- How much future context is needed to preserve names, tonal distinctions and language-switch boundaries?
- Can partial transcripts remain useful while making necessary corrections visible?
- How much latency comes from the model, and how much comes from buffering, turn detection and the call channel?
Transcribing Code-Switched Speech
Many multilingual evaluations assign one language to an utterance. Real conversations do not always respect that boundary. A Nigerian caller may speak accented English, move into Pidgin, and use an Igbo or Yorùbá phrase without pausing. Pidgin has its own vocabulary and grammar; treating it simply as English with pronunciation differences misses part of the problem.
A model can appear fluent while normalizing the caller’s words into something they did not say. Translation can hide recognition errors, and removing diacritics can hide distinctions that matter in tonal languages. A correct-looking English transcript is not sufficient evidence of understanding the original speech.
NaijaVoices demonstrates the value of language-specific data through adaptation experiments on Igbo, Hausa and Yorùbá. The remaining challenge is coverage of spontaneous conversation: mixed-language speech, regional variation, local names, business vocabulary and telephone channels.
Nigeria provides a concrete setting for these questions. The wider African challenge requires language-specific evidence rather than assuming that improvements in one market transfer to another. Shared architectures can help; local pronunciation, grammar and conversational conventions still need their own data and evaluation.
Speech Recognition with Background Speakers
Put a voice agent on a call from a market, a roadside shop or public transport, and every nearby voice can become part of the transcript. A station announcement may trigger a reply. Someone beside the caller may interrupt the agent without ever addressing it.
Why can’t noise suppression solve this? Generator hum and wind are different from another person speaking. Both the caller and the interfering speaker contain valid speech. A system needs some basis for deciding which one belongs to the conversation; simply keeping the loudest voice is not enough.
Speaker-conditioned extraction introduces another difficulty: the reference speech may itself be noisy or overlapping. The caller can change microphones or hand the phone to someone else. Aggressive extraction can also remove short acknowledgments or distort the acoustic detail needed to recognize a word.
- Caller selection: distinguish the intended speaker from nearby speech, even when the caller is quieter.
- Speech preservation: retain short words, names and tone-sensitive distinctions after enhancement.
- Interaction: distinguish genuine interruptions from background speech and the agent’s own playback.
- Latency: account for the audio context needed by the extractor, not only its processing time.
Better-sounding audio does not necessarily produce a better transcript. Recognition errors and false conversational turns are more useful tests of this problem than audio quality alone.
Speech Generation with Local Pronunciation
The earlier agenda questioned whether expression tags are enough to make speech adaptive. That question remains. A tag can request a style, but it does not carry all the information in the caller’s rhythm, emphasis, hesitation or previous turns.
Qwen3-TTS advances controllable and streaming synthesis. But multilingual generation is not the same as reliable local pronunciation. An English sentence can contain a Yorùbá name, an Igbo place name and a naira amount. The voice has to preserve all three without anglicizing the names or changing the number.
This is particularly important for repayment reminders, delivery confirmations and appointment calls. An omitted date or repeated amount can change the message. Naturalness scores alone do not capture that failure.
- Can synthesis move between languages within a sentence while preserving pronunciation and speaker identity?
- Does conditioning on previous audio improve delivery beyond text-only style instructions?
- Can the model remain expressive without omitting, repeating or changing the requested words?
Conversational appropriateness also needs local judgment. A pause or expression does not have one universal meaning across Africa, and vocal cues are not a reliable readout of someone’s internal emotional state. The relevant question is whether the response sounds clear, respectful and appropriate in that conversation.
Duplex Conversational Models for African Languages
Silence is an imperfect signal that a person has finished. A caller may pause to remember a landmark, correct an order while the agent is replying, or give a short acknowledgment that should not stop the response. A system that treats every sound as a new turn will interrupt; one that ignores speech during playback will miss corrections.
Moshi models user and system speech in parallel streams, giving duplex conversation a concrete research foundation. Qwen3-Omni connects understanding and speech generation through a Thinker–Talker architecture. Unified audio interaction creates opportunities to retain cues that disappear when speech is reduced to text.
But concurrent audio streams do not automatically produce good turn-taking. The model still needs to distinguish a correction from agreement, a background voice from the caller, and a language-switch pause from the end of a sentence. Speech-to-speech and modular voice systems should be judged on these behaviors, alongside response relevance and latency.
Better Evaluation and Post-training of Speech Models
A transcript can have a low word error rate and still get the most important detail wrong. Confusing ₦15,000 with ₦50,000, missing “not,” or changing one digit in a phone number can invalidate a whole interaction. Conversely, a harmless spelling variation can raise word error rate without changing the customer’s request.
Voice-agent evaluation therefore needs several layers: recognition of the original speech, preservation of critical details, appropriate conversation, and completion of the requested task. Consider the distinction between “I have paid” and a verified payment. Understanding the sentence does not establish that the transaction succeeded.
Nigerian business workflows make these distinctions concrete: qualifying a solar-installation enquiry, confirming a pay-on-delivery address given through landmarks, booking a clinic visit, or escalating an unresolved transfer complaint. These are useful evaluation settings because errors have a visible consequence beyond an imperfect transcript.
- Recognition fidelity: preserve names, amounts, dates, negation and language switches.
- Generation fidelity: avoid omitted words, repetitions and unintended speaker changes.
- Interaction: handle corrections, interruptions and uncertainty appropriately.
- Action correctness: distinguish a customer’s statement from verified system information and avoid duplicate actions.
Post-training for speech needs objectives that address these failures. Results should also be reported by language, speaking style and channel; an aggregate benchmark can hide poor performance for a particular group. The broader research question is how to make speech models fluent, context-aware and dependable across the conversations African customers actually have.
