Benchmarks

NVIDIA: Llama 3.1 Nemotron 70B Instruct

nvidia/llama-3.1-nemotron-70b-instruct

NVIDIA's Llama 3.1 Nemotron 70B is a language model designed for generating precise and useful responses. Leveraging [Llama 3.1 70B](/models/meta-llama/llama-3.1-70b-instruct) architecture and Reinforcement Learning from Human Feedback (RLHF), it excels...

Prompt $/1M
$1.200
Completion $/1M
$1.200
Context
131,072 tok
Best provider
26 tok/s · DeepInfra

Specs

Model id
nvidia/llama-3.1-nemotron-70b-instruct
Display name
NVIDIA: Llama 3.1 Nemotron 70B Instruct
Modality
text->text
Tokenizer
Llama3
Instruct type
llama3
Context length
131,072
Max completion
16,384
Prompt $/1M
$1.200
Completion $/1M
$1.200
Free tier
no

Capabilities

  • Function callingno
  • JSON modeno
  • System promptyes
  • Visionno
  • Moderatedno

Benchmark scores

Each axis is one benchmark, normalized to the best model in the field.

No benchmark data for this model.

Speed by provider

Providertok/s
DeepInfrafastest26
nvidiafastest
NVIDIA: Llama 3.1 Nemotron 70B Instruct - Benchmarks, Pricing and Speed | TryAii