NVIDIA: Llama 3.1 Nemotron Ultra 253B v1
nvidia/llama-3.1-nemotron-ultra-253b-v1
Llama-3.1-Nemotron-Ultra-253B-v1 is a large language model (LLM) optimized for advanced reasoning, human-interactive chat, retrieval-augmented generation (RAG), and tool-calling tasks. Derived from Meta’s Llama-3.1-405B-Instruct, it has been significantly customized using Neural...
Prompt $/1M
$0.600
Completion $/1M
$1.800
Context
131,072 tok
Best provider
39 tok/s · Nebius
Specs
- Model id
- nvidia/llama-3.1-nemotron-ultra-253b-v1
- Display name
- NVIDIA: Llama 3.1 Nemotron Ultra 253B v1
- Modality
- text->text
- Tokenizer
- Llama3
- Instruct type
- —
- Context length
- 131,072
- Max completion
- —
- Prompt $/1M
- $0.600
- Completion $/1M
- $1.800
- Free tier
- no
Capabilities
- Function callingno
- JSON modeno
- System promptno
- Visionno
- Moderatedno
Benchmark scores
Each axis is one benchmark, normalized to the best model in the field.
No benchmark data for this model.
Speed by provider
| Provider | tok/s |
|---|---|
| Nebiusfastest | 39 |