Our Research Agenda

    Pushing the frontier of capabilities in Speech models

    By Munachi Ernest-Eze

    Streaming, Contextual and Conversational Speech Models for Africa

    Our earlier research agenda explored streaming transcription, efficient audio representations, adaptive speech generation and duplex conversation. Recent work makes several of these directions more concrete. But faster speech models do not automatically understand African languages, and expressive voices do not automatically produce useful conversations.

    Consider a customer saying, “I don transfer the money, but e never reflect.” The model needs to preserve the Pidgin, understand the payment complaint, and distinguish it from confirmation that a payment succeeded. Add another person speaking nearby, a local name and a correction to an amount, and the problem becomes much larger than transcribing a clean sentence.

    Realtime Multilingual Transcription

    Why does streaming remain difficult? A model has to produce text before it has heard the full utterance. More future audio can resolve ambiguity, but waiting for it adds delay. Emitting text earlier can introduce revisions, especially around short words, names and language switches.

    Voxtral Realtime addresses this through a causal audio encoder and aligned audio-text streams during training. Qwen3-ASR offers smaller recognition models alongside a larger variant. These are useful advances beyond treating every incoming chunk as another offline transcription.

    However, a fast first token does not mean the transcript is already reliable. A caller may correct a delivery location halfway through a sentence or switch into Yorùbá after giving an amount in English. The important tradeoff is between delay and stable meaning, not just delay and the first visible word.

    • How much future context is needed to preserve names, tonal distinctions and language-switch boundaries?
    • Can partial transcripts remain useful while making necessary corrections visible?
    • How much latency comes from the model, and how much comes from buffering, turn detection and the call channel?

    Transcribing Code-Switched Speech

    Many multilingual evaluations assign one language to an utterance. Real conversations do not always respect that boundary. A Nigerian caller may speak accented English, move into Pidgin, and use an Igbo or Yorùbá phrase without pausing. Pidgin has its own vocabulary and grammar; treating it simply as English with pronunciation differences misses part of the problem.

    A model can appear fluent while normalizing the caller’s words into something they did not say. Translation can hide recognition errors, and removing diacritics can hide distinctions that matter in tonal languages. A correct-looking English transcript is not sufficient evidence of understanding the original speech.

    NaijaVoices demonstrates the value of language-specific data through adaptation experiments on Igbo, Hausa and Yorùbá. The remaining challenge is coverage of spontaneous conversation: mixed-language speech, regional variation, local names, business vocabulary and telephone channels.

    Nigeria provides a concrete setting for these questions. The wider African challenge requires language-specific evidence rather than assuming that improvements in one market transfer to another. Shared architectures can help; local pronunciation, grammar and conversational conventions still need their own data and evaluation.

    Speech Recognition with Background Speakers

    Put a voice agent on a call from a market, a roadside shop or public transport, and every nearby voice can become part of the transcript. A station announcement may trigger a reply. Someone beside the caller may interrupt the agent without ever addressing it.

    Why can’t noise suppression solve this? Generator hum and wind are different from another person speaking. Both the caller and the interfering speaker contain valid speech. A system needs some basis for deciding which one belongs to the conversation; simply keeping the loudest voice is not enough.

    Speaker-conditioned extraction introduces another difficulty: the reference speech may itself be noisy or overlapping. The caller can change microphones or hand the phone to someone else. Aggressive extraction can also remove short acknowledgments or distort the acoustic detail needed to recognize a word.

    • Caller selection: distinguish the intended speaker from nearby speech, even when the caller is quieter.
    • Speech preservation: retain short words, names and tone-sensitive distinctions after enhancement.
    • Interaction: distinguish genuine interruptions from background speech and the agent’s own playback.
    • Latency: account for the audio context needed by the extractor, not only its processing time.

    Better-sounding audio does not necessarily produce a better transcript. Recognition errors and false conversational turns are more useful tests of this problem than audio quality alone.

    Speech Generation with Local Pronunciation

    The earlier agenda questioned whether expression tags are enough to make speech adaptive. That question remains. A tag can request a style, but it does not carry all the information in the caller’s rhythm, emphasis, hesitation or previous turns.

    Qwen3-TTS advances controllable and streaming synthesis. But multilingual generation is not the same as reliable local pronunciation. An English sentence can contain a Yorùbá name, an Igbo place name and a naira amount. The voice has to preserve all three without anglicizing the names or changing the number.

    This is particularly important for repayment reminders, delivery confirmations and appointment calls. An omitted date or repeated amount can change the message. Naturalness scores alone do not capture that failure.

    • Can synthesis move between languages within a sentence while preserving pronunciation and speaker identity?
    • Does conditioning on previous audio improve delivery beyond text-only style instructions?
    • Can the model remain expressive without omitting, repeating or changing the requested words?

    Conversational appropriateness also needs local judgment. A pause or expression does not have one universal meaning across Africa, and vocal cues are not a reliable readout of someone’s internal emotional state. The relevant question is whether the response sounds clear, respectful and appropriate in that conversation.

    Duplex Conversational Models for African Languages

    Silence is an imperfect signal that a person has finished. A caller may pause to remember a landmark, correct an order while the agent is replying, or give a short acknowledgment that should not stop the response. A system that treats every sound as a new turn will interrupt; one that ignores speech during playback will miss corrections.

    Moshi models user and system speech in parallel streams, giving duplex conversation a concrete research foundation. Qwen3-Omni connects understanding and speech generation through a Thinker–Talker architecture. Unified audio interaction creates opportunities to retain cues that disappear when speech is reduced to text.

    But concurrent audio streams do not automatically produce good turn-taking. The model still needs to distinguish a correction from agreement, a background voice from the caller, and a language-switch pause from the end of a sentence. Speech-to-speech and modular voice systems should be judged on these behaviors, alongside response relevance and latency.

    Better Evaluation and Post-training of Speech Models

    A transcript can have a low word error rate and still get the most important detail wrong. Confusing ₦15,000 with ₦50,000, missing “not,” or changing one digit in a phone number can invalidate a whole interaction. Conversely, a harmless spelling variation can raise word error rate without changing the customer’s request.

    Voice-agent evaluation therefore needs several layers: recognition of the original speech, preservation of critical details, appropriate conversation, and completion of the requested task. Consider the distinction between “I have paid” and a verified payment. Understanding the sentence does not establish that the transaction succeeded.

    Nigerian business workflows make these distinctions concrete: qualifying a solar-installation enquiry, confirming a pay-on-delivery address given through landmarks, booking a clinic visit, or escalating an unresolved transfer complaint. These are useful evaluation settings because errors have a visible consequence beyond an imperfect transcript.

    • Recognition fidelity: preserve names, amounts, dates, negation and language switches.
    • Generation fidelity: avoid omitted words, repetitions and unintended speaker changes.
    • Interaction: handle corrections, interruptions and uncertainty appropriately.
    • Action correctness: distinguish a customer’s statement from verified system information and avoid duplicate actions.

    Post-training for speech needs objectives that address these failures. Results should also be reported by language, speaking style and channel; an aggregate benchmark can hide poor performance for a particular group. The broader research question is how to make speech models fluent, context-aware and dependable across the conversations African customers actually have.

    •By Munachi Ernest-Eze

    Streaming, Fast, Multilingual Speech to Text models

    Whisper was released in Dec'22, but even then, most speech to text models of today suffer from the same challenges as whisper because they either use whisper as encoder (ex - Voxtral) or use a similar architecture. Specifically STT model suffer from the following challenges:

    • Realtime Transcription: They aren't designed to transcribe audio realtime (i.e, every 500ms) fast.
    • They are either slow (/big) & multi-lingual or fast (/small) but error-prone & english only, i.e, monlingual.
    • They use too many audio tokens/frames. Ex - Whisper models operate at 50 audio tokens/s.

    Realtime Transcription - Transcribing 500ms audio fast

    You may ask why does realtime transcription even matter? The answer is that this results in a much better UX experience - user can see what their speech is being transcribed into, get realtime feedback, and then correct it mid sentence.

    For example, see the attached demo of Wisper Flow, it only transcribes your speech once you finish the whole sentence - so no realtime feedback.

    Many Language Learning apps like Stimuler, MySivi have live AI audio calls. The user isn't expected to make mistakes in the language they're learning - and hence realtime transcriptions are even more critical, for users to correct their mistakes before AI incorrectly interprets & responds to them.

    Why can't current STT models transcribe in realtime? Whisper uses a fixed input length of 30s, irrespective of your audio length. At 50 tokens / sec, each input audio is padded to a whopping 1500 tokens. On an A100, this would mean:

    • ~40ms prefill time for whisper-large-v3 / whisper-large-v3-turbo. Furthermore, prefills this large are compute bound and can't be parallelized. [1]
    • Significantly slow decode time because of heavy KV Cache read from HBM. Whisper-large-v3-turbo uses a small decoder to speed up decode time, but it doesn't resolve the fundamental issue of slow KVcache reads.

    Voxtral tried making input lengths dynamic, but that resulted in loss of performance.

    Voxtral tried removing padding but that resulted in a slight loss in performance

    Intuitively this makes sense, but should have been handled differently. Padded tokens seem to be acting as register tokens offering more compute for transcription for small input sequences leading to better WER. We should remove padding completely and allow users to add custom number of "registers" to improve accuracy.

    Small & Capable Multilingual Transcription

    Speech to Text models suffer from an expected tradeoff. You can either make them big and powerful but slow, or you can make them small and monolingual & fast with higher error rate.

    LLMs use MOEs to solve for this tradeoff. Ideal, speech to transcription models should have latencies of a 100M model but capabilities of a 3B model.

    Fewer Tokens per frame

    Most speech transcription models operate at 50Hz (audio tokens per second) which leads to very large sequence lengths and both slower prefill & decode times during inference.

    Many other works in speech have shown that 12.5Hz is capable enough to model audio (ex - Moshi). Using 12.5Hz for audio transcription instead of 50Hz would allow us to shrink sequence length by 4x and will reduce KVCache size & lead to faster prefills & decode.

    Speech Generation with Adaptive Expressiveness

    AI generated speech that can adapt to user's speech & tone.

    Speech models have become really expressive by using LLM backbones. Example - ElevenLabs V3, Sesame's CSM-1B but in most of the approaches - the expressiveness is robotic and non-adaptive.

    ElevelLabs V3, and Canopy Labs's Orpheus and even nari-labs/Dia use tags like <laugh>, <cry> to make the model's speech expressive. This approach is like putting a duct-tape on a broken pipe. It suffers from these fundamental issues:

    • Model's expressiveness isn't adaptive: These tags come from LLM which only sees user's speech's transcription and not the actual speech. User could say "Stop" politely or in anger, or in disappointment which LLM never sees.
    • Tags aren't granular: These tags don't tell the model how much to laugh? whether to keep on laughing or just laugh a little before continuing, whether to laugh sarcastically or genuinely.

    Sesame's CSM-1B takes user's audio speech into account to generate the corresponding model's speech based on some text. Although other than vibe-checks the effect of this hasn't been measured. This is not as easy as it sounds - because most of the conversational data to train for this comes from podcast which don't really contain a lot of edge cases - like of user's crying or being emotional, or angry.

    To build speech generation with adaptive expressiveness, we need:

    • Evals: that can measure how well the model's speech adapts to user's change in tone & expressiveness.
    • Data: Synthetic or otherwise conversational data with variety of speakers emotions.

    Duplex Conversational Models

    Conversational Models that can speak & listen at the same time

    Currently, a Voice Activity Detection (VAD) module is used to detect when to trigger the speech model in both Voice Agents & realtime models. A normal VAD operates on the audio waveform and decides if there's silence, and a semantic VAD triggers a small LLM / classifier to decide whether the user has finished their sentence based on the user's speech's text.

    Because of this models can either speak or listen, but can't do both simultaneously leading to abrupt & unnatural interruptions. We believe no matter how expressive or adaptive the AI voices become, this is one of the great tells that you're talking to an AI rather than a human.

    Current conversational models can't do the following, which humans do casually & frequently:

    • interrupt user intelligently - knowing that they haven't finished their sentence.
    • when user interrupts you - give a slight nudge acknowledging their interruption, but still choose to finish your sentence. once you finish your whole sentence, then ask user about what they were saying earlier.

    Human's lag for speaking and listening is also only ~100ms. So we need to build fast & efficient two-way duplex conversational models that can speak and listen at the same time, where the model receives user's input every 100ms irrespective of whether it's speaking or not, and decides to speak, nudge, or stay silent.

    Modelling long-form speech

    Dub a 2hr long film in one shot

    Model NameTokens / hrDescription
    OpenAI/Whisper180KAudio Encoder, samples audio @ ~50Hz
    Mistral/Voxtral45KAudioLM (audio input-only), uses whisper encoder but downsamples audio by 4x to 12.5Hz
    kyutai/Moshi45KAudioLM (audio input+output) uses Mimi codec which samples 32 RVQ tokens @ 12.5Hz
    CanopyLabs/Orpheus315KUses Snac tokenizer with hierarchical 7 tokens @ 12.5Hz; flattens all tokens per frame.

    This is kind-of okay for Audio-understanding tasks, since even LLMs are now routinely trained on 128K input context length.

    Although, this is really troubling for complex tasks on long-form audio; like audio editing, or dubbing a long movie in one shot - since even LLM struggle to generate more than 4-8K output tokens despite significant demand for this in the past few years.

    To solve long-form audio gen tasks, we either need to get #tokens down, or design new frameworks that can solve these tasks iteratively in an agentic fashion.

    Better Post-training of speech models

    Post-training speech models to add audio-editing & Multi-task capabilities

    Current post-training of speech models is significantly behind compared to LLMs. Most of the leading models have similar failure artifacts as repeating a word or not following instructions correctly - remember my grandma will die if you don't give me a valid json?

    Gemini-2.5-Pro-TTS fails on a simple prompt. It not only doesn't speak the full sentence but also mixes speaker 1 and 2.

    Dia-TTS fails to say the correct sentence or switch speakers.

    And this is just for TTS. To build speech models that can do complex tasks like multi-character dubbing, voice cloning, TTS, STT in the same model we need to have robust post-training to prevent from similar failures.

    Furthermore, speech models are still fragmented across various tasks - TTS, VoiceCloning, STT & do not support complex tasks at all. There have been efforts to make these models multi-task, for example Moshi can do both TTS, STT, Audio QnA in the same model.

    To push the frontier of post-training in speech models, we need to:

    • Create better evals for Instruction following in multi-task speech models.
    • Extend speech models to complex speech tasks like audio-editing, multi-character audio dubbing.

    Appendix

    [1] Algorithmic Intensity for A100 is FLOPs / Memory Bandwidth = 312e12 FLOPs/ 2,039 GB/s = 153. Thus, prefill with more than 153 tokens would be compute bound as per roofline analysis.

    Voice AI for your customer operations.

    Discuss outbound sales, customer support and intake in the languages your customers speak.

    Talk to our team