Qwen: Qwen3 VL 32B Instruct
qwen/qwen3-vl-32b-instruct
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...
Prompt $/1M
$0.104
Completion $/1M
$0.416
Context
262,144 tok
Best provider
96 tok/s · Artificial Analysis
Specs
- Model id
- qwen/qwen3-vl-32b-instruct
- Display name
- Qwen: Qwen3 VL 32B Instruct
- Modality
- text+image->text
- Tokenizer
- Qwen
- Instruct type
- —
- Context length
- 262,144
- Max completion
- 32,768
- Prompt $/1M
- $0.104
- Completion $/1M
- $0.416
- Free tier
- no
Capabilities
- Function callingno
- JSON modeno
- System promptno
- Visionyes
- Moderatedno
Benchmark scores
Each axis is one benchmark, normalized to the best model in the field.
No benchmark data for this model.
Speed by provider
| Provider | tok/s |
|---|---|
| Artificial Analysis | 96 |
| Alibabafastest | 37 |