NVIDIA: Llama 3.1 Nemotron 70B Instruct
nvidia/llama-3.1-nemotron-70b-instruct
NVIDIA's Llama 3.1 Nemotron 70B is a language model designed for generating precise and useful responses. Leveraging [Llama 3.1 70B](/models/meta-llama/llama-3.1-70b-instruct) architecture and Reinforcement Learning from Human Feedback (RLHF), it excels...
Prompt $/1M
$1.200
Completion $/1M
$1.200
Context
131,072 tok
Best provider
26 tok/s · DeepInfra
Specs
- Model id
- nvidia/llama-3.1-nemotron-70b-instruct
- Display name
- NVIDIA: Llama 3.1 Nemotron 70B Instruct
- Modality
- text->text
- Tokenizer
- Llama3
- Instruct type
- llama3
- Context length
- 131,072
- Max completion
- 16,384
- Prompt $/1M
- $1.200
- Completion $/1M
- $1.200
- Free tier
- no
Capabilities
- Function callingno
- JSON modeno
- System promptyes
- Visionno
- Moderatedno
Benchmark scores
Each axis is one benchmark, normalized to the best model in the field.
No benchmark data for this model.
Speed by provider
| Provider | tok/s |
|---|---|
| DeepInfrafastest | 26 |
| nvidiafastest | — |