Benchmarks

Xiaomi: MiMo-V2-Omni

xiaomi/mimo-v2-omni

MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step...

Prompt $/1M
$0.400
Completion $/1M
$2.000
Context
262,144 tok
Best provider
16 tok/s · Xiaomi

Specs

Model id
xiaomi/mimo-v2-omni
Display name
Xiaomi: MiMo-V2-Omni
Modality
text+image+audio+video->text
Tokenizer
Other
Instruct type
Context length
262,144
Max completion
65,536
Prompt $/1M
$0.400
Completion $/1M
$2.000
Free tier
no

Capabilities

  • Function callingno
  • JSON modeno
  • System promptno
  • Visionyes
  • Moderatedno

Benchmark scores

Each axis is one benchmark, normalized to the best model in the field.

No benchmark data for this model.

Speed by provider

Providertok/s
Xiaomifastest16
Xiaomi: MiMo-V2-Omni - Benchmarks, Pricing and Speed | TryAii