DeepSeek-V4.1-Flash: GPU memory and hosting cost
Open weights · 763.21 billion parameters · Mixture of experts · 1,048,576-token context · licence mit
No cloud sells this model per token: to use it you run it yourself, on GPU instances you rent.
Run it yourself
It needs 1908 GB of GPU memory at fp16, the precision it ships in.
| Cloud | Cheapest instance that fits | GPUs | An hour | A month, around the clock |
|---|---|---|---|---|
| AWS | p6-b300.48xlarge | 8 × B300 268 GB | $142.42 | $103,964 |
More from DeepSeek
Similar size, open weights
- Llama 4 Maverick 17B
- Llama 3.1 405B
- GLM-5.3
- GLM-5.2
- Qwen3-Coder-480B-A35B-Instruct-FP8
- NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
List prices collected 6 October 2026, refreshed every three hours. No negotiated rates or taxes.