Xiaomi: MiMo-V2-Omni
xiaomi/mimo-v2-omni
MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step...
Prompt $/1M
$0.400
Completion $/1M
$2.000
Context
262,144 tok
Best provider
16 tok/s · Xiaomi
Specs
- Model id
- xiaomi/mimo-v2-omni
- Display name
- Xiaomi: MiMo-V2-Omni
- Modality
- text+image+audio+video->text
- Tokenizer
- Other
- Instruct type
- —
- Context length
- 262,144
- Max completion
- 65,536
- Prompt $/1M
- $0.400
- Completion $/1M
- $2.000
- Free tier
- no
Capabilities
- Function callingno
- JSON modeno
- System promptno
- Visionyes
- Moderatedno
Benchmark scores
Each axis is one benchmark, normalized to the best model in the field.
No benchmark data for this model.
Speed by provider
| Provider | tok/s |
|---|---|
| Xiaomifastest | 16 |