Ling-3.0-flash
inclusionai/ling-3.0-flash
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...
Prompt $/1M
$0.075
Completion $/1M
$0.220
Context
131,072 tok
Best provider
322 tok/s · DeepInfra
Specs
- Model id
- inclusionai/ling-3.0-flash
- Display name
- Ling-3.0-flash
- Modality
- text->text
- Tokenizer
- Other
- Instruct type
- —
- Context length
- 131,072
- Max completion
- 16,384
- Prompt $/1M
- $0.075
- Completion $/1M
- $0.220
- Free tier
- no
Capabilities
- Function callingno
- JSON modeno
- System promptno
- Visionno
- Moderatedno
Benchmark scores
Each axis is one benchmark, normalized to the best model in the field.
No benchmark data for this model.
Speed by provider
| Provider | tok/s |
|---|---|
| DeepInfrafastest | 322 |
| Artificial Analysis | 316 |