What Kyutai does
Kyutai is a Paris-based non-profit open-science AI lab focused on building and democratizing artificial general intelligence through open research. Its work centers on multimodal, low-latency, real-time model capabilities—especially speech-native systems—released through scientific papers, open-source code/weights, and public demos.
More
Kyutai’s product portfolio is organized around “speech-native” and “cascaded voice” building blocks that can be combined into voice agents: - Moshi is described by Kyutai as a real-time, speech-native dialogue system (audio-native interaction). - Hibiki is Kyutai’s speech-to-speech translation model designed for simultaneous translation while preserving the speaker’s voice (initially demonstrated French→English). - Kyutai STT (speech-to-text) is a streaming architecture optimized for interactive applications, supporting English/French (and an English-only larger model). - Kyutai TTS includes an open-sourced “Pocket TTS” (100M parameters) and “TTS 1.6B” described as streaming and suitable for low-latency voice assistants and server use. - Unmute is Kyutai’s modular system that lets an arbitrary text LLM listen and speak using Kyutai’s real-time STT and TTS; Kyutai emphasizes semantic VAD to reduce interruption, and fast end-to-end latency via streaming generation. Beyond speech, Kyutai also pursues multimodal vision and world-model research (e.g., CASA and MIRA for video/world modeling), aiming to connect modalities efficiently and to release related research artifacts. Business model / positioning: Kyutai states it is funded by its founding donors (notably Iliad Group, CMA CGM Group, and Schmidt Sciences / Schmidt Futures). Unlike a typical vendor, Kyutai’s public-facing “value proposition” is open research outputs (weights, code, demos, and tutorials) that other developers and researchers can build on for voice and multimodal applications.