Hibiki-Zero v1.0
Hibiki-Zero is an advanced model for simultaneous speech-to-speech translation that eliminates the need for word-level aligned data, simplifying the training pipeline and enabling seamless scaling to diverse languages. It achieves state-of-the-art performance in translation accuracy, latency, voice transfer, and naturalness across multiple tasks.
Simultaneous speech-to-speech translation without word-level aligned data
State-of-the-art performance in translation accuracy and latency
Preserves speaker identity and speech naturalness