LLM Opex Every way to run an LLM, every cloud

Where the prices come from, and how fresh they are

Every figure on the site is a public list price, collected automatically from each cloud's own price list or documentation every three hours. No negotiated rates, commitments or taxes.

Prices held, by cloud

CloudPricesLast collectedAge
Alibaba Cloud1,3225 October 20262 hours ago
AWS19,9645 October 20263 hours ago
Azure36,0585 October 20263 hours ago
Core421045 October 20263 hours ago
Google Cloud3095 October 20262 hours ago
OCI1225 October 20263 hours ago

The public sources

AWS

  • GPU instances: https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/AmazonEC2/current/{region}/index.csv
  • Bedrock token prices: https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/{AmazonBedrock,AmazonBedrockService,AmazonBedrockFoundationModels}/current/index.json
  • freshness overlay: https://b0.p.awsstatic.com/pricing/2.0/meteredUnitMaps/bedrock/USD/current/bedrock.json

Azure

  • GPU VMs and token meters: https://prices.azure.com/api/retail/prices

OCI

  • GPU shapes and Generative AI meters: https://apexapps.oracle.com/pls/apex/cetools/api/v1/products/
  • models offered, and in which mode: https://docs.oracle.com/en-us/iaas/Content/generative-ai/pretrained-models.htm
  • on-demand vs dedicated, per region: https://docs.oracle.com/en-us/iaas/Content/generative-ai/model-endpoint-regions.htm
  • dedicated cluster AI unit counts: https://docs.oracle.com/en-us/iaas/Content/generative-ai/hardware-unit-shapes-by-region.htm

Core42

  • per-token prices: https://www.core42.ai/compass/documentation/compass-model-pricing
  • models, tasks and the region each runs in: https://www.core42.ai/compass/documentation/compass-models
  • deprecation and retirement dates: https://www.core42.ai/compass/documentation/compass-model-deprecations-and-retirements

Alibaba Cloud

  • per-token prices (Model Studio): https://www.alibabacloud.com/help/en/model-studio/model-pricing
  • GPU instance prices (ECS OpenAPI): https://www.alibabacloud.com/help/en/ecs/developer-reference/api-ecs-2014-05-26-describeprice
  • GPU instance specifications: https://www.alibabacloud.com/help/en/ecs/user-guide/gpu-accelerated-compute-optimized-and-vgpu-accelerated-instance-families-1

Google Cloud

  • per-token prices (Vertex AI): https://cloud.google.com/vertex-ai/generative-ai/pricing
  • GPU machine prices (Compute Engine): https://cloud.google.com/products/compute/pricing/accelerator-optimized
  • GPU machine specifications: https://cloud.google.com/compute/docs/gpus
  • internet egress (Premium Tier): https://cloud.google.com/vpc/network-pricing