Qwen: Qwen3 VL 8B Instruct
qwen/qwen3-vl-8b-instruct
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
Prompt $/1M
$0.117
Completion $/1M
$0.455
Context
256,000 tok
Best provider
141 tok/s · Artificial Analysis
Specs
- Model id
- qwen/qwen3-vl-8b-instruct
- Display name
- Qwen: Qwen3 VL 8B Instruct
- Modality
- text+image->text
- Tokenizer
- Qwen3
- Instruct type
- —
- Context length
- 256,000
- Max completion
- 32,768
- Prompt $/1M
- $0.117
- Completion $/1M
- $0.455
- Free tier
- no
Capabilities
- Function callingno
- JSON modeno
- System promptno
- Visionyes
- Moderatedno
Benchmark scores
Each axis is one benchmark, normalized to the best model in the field.
No benchmark data for this model.
Speed by provider
| Provider | tok/s |
|---|---|
| Artificial Analysis | 141 |
| Alibabafastest | 65 |
| AtlasCloud | 40 |
| Parasail | 37 |
| Novita | 30 |