We build speech foundation models.
We started with echo, our family of speech-native LLMs — language models that listen to audio directly and reason over it, instead of transcribing first and thinking later. echo is trained on English and Arabic across a wide range of dialects, so it holds up on the accents and code-switching that real speech actually contains, not just the clean read-aloud audio most models are tuned for.
The next phase is Conversational Speech Models, our take on a new kind of TTS. Rather than reading a sentence back in a fixed voice, these models are built for real conversation — speech that carries the timing, tone and turn-taking of an actual dialogue.
Everything here is Arabic-first and built for production, from the data and the encoders through to low-latency serving over telephony and far-field audio.