Benchmarks

NVIDIA: Nemotron 3 Ultra

nvidia/nemotron-3-ultra-550b-a55b

Nemotron 3 Ultra is an open-weight heavyweight from NVIDIA — 550B parameters with 55B active per token, ranking #90 of 235 on the TryAii Score. It's strongest on head-to-head chat (#32 on Chatbot Arena Elo). API pricing starts at $0.625/$3.12 per million input/output tokens, though similar scores now sell for about $1.2 per million output tokens, with 6 independent hosts serving it. It's among the fastest models we track, at roughly 28-104 tokens/sec.

Prompt $/1M
$0.625
Completion $/1M
$3.125
Context
262,144 tok
Best provider
137 tok/s · Nebius

Specs

Model id
nvidia/nemotron-3-ultra-550b-a55b
Display name
NVIDIA: Nemotron 3 Ultra
Modality
text->text
Tokenizer
Other
Instruct type
Context length
262,144
Max completion
32,768
Prompt $/1M
$0.625
Completion $/1M
$3.125
Free tier
no

Capabilities

  • Function callingno
  • JSON modeno
  • System promptno
  • Visionno
  • Moderatedno

Benchmark scores

Each axis is one benchmark, normalized to the best model in the field.

No benchmark data for this model.

Speed by provider

Providertok/s
Nebiusfastest137
Crusoe111
BaseTen83
Together70
Venice14
DeepInfra10
nvidiafastest