Xiaomi
MiMo-V2-Omni v2.0
MiMo-V2-Omni is Xiaomi's latest multimodal foundation model designed to process and understand images, video, and audio simultaneously. It enables agents to perceive and act in the real world by integrating these modalities into a unified perceptual stream, facilitating real-time, complex decision-making across various applications.
MiMo-V2-TTS vV2
MiMo-V2-TTS is Xiaomi's self-developed large-scale speech synthesis model, designed to provide AI agents with expressive, natural-sounding voices. It offers contextual emotion awareness, universal style adaptability, and real-time, seamless interaction, enabling agents to communicate with warmth and authenticity.
MiMo-V2.5-Pro-UltraSpeed vV2.5-Pro-UltraSpeed
MiMo-V2.5-Pro-UltraSpeed is an advanced AI language model developed by Xiaomi, capable of generating up to 1,200 tokens per second on a 1-trillion-parameter model. This performance is achieved through a deep collaboration with TileRT, focusing on extreme model-system co-design to optimize inference speed on commodity GPUs.
MiMo Code vV0.1.0
MiMo Code is an open-source AI coding assistant developed by Xiaomi, designed to enhance developer productivity by providing intelligent coding support directly within the terminal. It integrates Xiaomi's MiMo-V2.5 model and supports major large models such as DeepSeek and Kimi, offering a persistent memory system and advanced features for efficient coding tasks.