Benchmarks

NVIDIA: Llama 3.1 Nemotron Ultra 253B v1

nvidia/llama-3.1-nemotron-ultra-253b-v1

Llama-3.1-Nemotron-Ultra-253B-v1 is a large language model (LLM) optimized for advanced reasoning, human-interactive chat, retrieval-augmented generation (RAG), and tool-calling tasks. Derived from Meta’s Llama-3.1-405B-Instruct, it has been significantly customized using Neural...

Prompt $/1M
$0.600
Completion $/1M
$1.800
Context
131,072 tok
Best provider
39 tok/s · Nebius

Specs

Model id
nvidia/llama-3.1-nemotron-ultra-253b-v1
Display name
NVIDIA: Llama 3.1 Nemotron Ultra 253B v1
Modality
text->text
Tokenizer
Llama3
Instruct type
Context length
131,072
Max completion
Prompt $/1M
$0.600
Completion $/1M
$1.800
Free tier
no

Capabilities

  • Function callingno
  • JSON modeno
  • System promptno
  • Visionno
  • Moderatedno

Benchmark scores

Each axis is one benchmark, normalized to the best model in the field.

No benchmark data for this model.

Speed by provider

Providertok/s
Nebiusfastest39
NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 - Benchmarks, Pricing and Speed | TryAii