Benchmarks

Meta: Llama 3.2 11B Vision Instruct

meta-llama/llama-3.2-11b-vision-instruct

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and...

Prompt $/1M
$0.345
Completion $/1M
$0.345
Context
131,072 tok
Best provider
70 tok/s · Artificial Analysis

Specs

Model id
meta-llama/llama-3.2-11b-vision-instruct
Display name
Meta: Llama 3.2 11B Vision Instruct
Modality
text+image->text
Tokenizer
Llama3
Instruct type
llama3
Context length
131,072
Max completion
16,384
Prompt $/1M
$0.345
Completion $/1M
$0.345
Free tier
no

Capabilities

  • Function callingno
  • JSON modeno
  • System promptyes
  • Visionyes
  • Moderatedno

Benchmark scores

Each axis is one benchmark, normalized to the best model in the field.

No benchmark data for this model.

Speed by provider

Providertok/s
Artificial Analysis70
DeepInfrafastest46
Meta: Llama 3.2 11B Vision Instruct - Benchmarks, Pricing and Speed | TryAii