What will my workload cost, and where is the break-even?
Enter the tokens you expect a day and the calculator prices the workload three ways on each cloud, side by side, and shows the volume at which the cheapest option changes.
The three ways to run a model
- Per token. You pay the cloud's list price for each million input and output tokens. Nothing to run; the cost follows your volume.
- Self-hosted. You rent GPU instances by the hour and run the model yourself. The cost is the instances the model needs, around the clock, whatever the volume; the price per token falls as you use them more.
- Dedicated cluster. The cloud runs the model for you on hardware reserved for you, for a fixed price per hour or per month.
How the figures are worked out
Per token: monthly input tokens × the input price, plus monthly output tokens × the output price. Self-hosted and dedicated: the hourly price × the hours in a month × the number of instances or units, where that number comes from the model's size, its precision and the throughput the hardware delivers. The break-even is the volume at which the fixed monthly cost equals the per-token cost. Every formula and source is on How it works.
Compare per-token prices across clouds or see which hardware a model fits on.