Example of it working in action: https://streamable.com/ueh3sj
Paper: https://arxiv.org/abs/2502.03382
Samples: https://hf.co/spaces/kyutai/hibiki-samples
Inference code: https://github.com/kyutai-labs/hibiki
Models: https://huggingface.co/kyutai
From kyutai on X: Meet Hibiki, our simultaneous speech-to-speech translation model, currently supporting FR to EN.
Hibiki produces spoken and text translations of the input speech in real-time, while preserving the speaker’s voice and optimally adapting its pace based on the semantic content of the source speech.
Based on objective and human evaluations, Hibiki outperforms previous systems for quality, naturalness and speaker similarity and approaches human interpreters.
https://x.com/kyutai_labs/status/1887495488997404732
Neil Zeghidour on X: https://x.com/neilzegh/status/1887498102455869775
Nice. That’s exactly what we need on our phones. And on video players. Just needs support for 20 other languages and be rolled out as open source.