Which GPU instance or cluster can actually hold it?
A model fits when its weights, plus the memory its attention cache needs for your context length and the requests you serve at once, are within the GPU memory of an instance or a cluster. Anything too small is left out; what fits is listed per cloud with its price.
Memory needed and the cheapest hardware that fits, by model
| Model | Parameters | Precision | GPU memory needed | Cheapest instance that fits |
|---|---|---|---|---|
| OpenAI gpt-oss 120B | 116.83 B | int4 | 73 GB | Azure Standard_NC24ads_A100_v4: 1 × A100, $3.67 an hour |
| OpenAI gpt-oss 20B | 20.91 B | int4 | 13.1 GB | AWS g4dn.xlarge: 1 × T4, $0.526 an hour |
| Llama 3.3 70B | 70.55 B | bf16 | 176.4 GB | Alibaba Cloud ecs.gn8v-2x.8xlarge: 2 × GPU H, $14.38 an hour |
| Llama 4 Maverick 17B | 401.58 B | bf16 | 1003.9 GB | AWS p5en.48xlarge: 8 × H200, $63.30 an hour |
| Llama 4 Scout 17B | 108.64 B | bf16 | 271.6 GB | Google Cloud a2-ultragpu-4g: 4 × A100, $20.28 an hour |
| Qwen3 32B | 32.76 B | bf16 | 81.9 GB | Google Cloud g4-standard-48: 1 × RTX PRO 6000, $4.50 an hour |
| DeepSeek V3.1 | 684.53 B | fp8 | 855.7 GB | AWS p5en.48xlarge: 8 × H200, $63.30 an hour |
| Llama 3.1 405B | 405.85 B | bf16 | 1014.6 GB | AWS p5en.48xlarge: 8 × H200, $63.30 an hour |
| Mistral Small | 23.57 B | bf16 | 58.9 GB | Azure Standard_NC24ads_A100_v4: 1 × A100, $3.67 an hour |
| Qwen3 235B A22B | 235.09 B | bf16 | 587.7 GB | AWS p4de.24xlarge: 8 × A100, $27.45 an hour |
| Llama 3.1 70B | 70.55 B | bf16 | 176.4 GB | Alibaba Cloud ecs.gn8v-2x.8xlarge: 2 × GPU H, $14.38 an hour |
| Llama 3.1 8B | 8.03 B | bf16 | 20.1 GB | Google Cloud g2-standard-4: 1 × L4, $0.7068 an hour |
| Phi-4 | 14.66 B | bf16 | 36.6 GB | AWS g6e.xlarge: 1 × L40S, $1.86 an hour |
| Mistral 7B | 7.25 B | bf16 | 18.1 GB | Google Cloud g2-standard-4: 1 × L4, $0.7068 an hour |