MiMo-V2.5
A 310B-param (15B active) MoE omnimodal model handling text, image, video & audio, with 1M context and strong agentic/coding capabilities.
Model details
MiMo-V2.5 is Xiaomi's flagship omnimodal AI model, built as a sparse Mixture-of-Experts architecture with 310B total parameters (15B activated) and a hybrid sliding-window/global attention design that supports context windows up to 1M tokens. It natively understands text, images, video, and audio through dedicated vision and audio encoders layered onto the MiMo-V2-Flash backbone, though its own outputs are text-only. MiMo-V2.5 is designed to excel at multimodal perception, long-context reasoning, and agentic workflows like coding and multi-step tool use.