How We Estimate Costs

Here's how we come up with the cost estimates

Our goal is to make prices from different providers as comparable as possible.

Keep in mind these are examples for common use cases, not quotes. Your actual costs will differ based on utilization, reserved instances, discounts, and other factors we haven't accounted for.

Assumptions

All estimates use the same baseline:

  • A month is 30 days (720 hours).
  • We use the region closest to North Virginia (USA) or Frankfurt (Germany).
  • We ignore temporary promotions and special discount programs.
  • We aim to record prices before VAT and sales tax. A provider that shows a tax-inclusive figure by default, based on where it thinks you are, may therefore quote more on its own site than we do.
  • All prices are shown in USD. Where a provider publishes in another currency we convert at the European Central Bank's daily reference rate. We currently track prices published in EUR, SEK, CHF, ISK and INR.

Reference configurations

Outside of GPUs, we compare providers against these reference configurations. Each one links to the full specification we price against.

Resource Configuration
VM Small 1 vCPU, 1 GB RAM
VM Medium 4 vCPU, 8 GB RAM
VM Large 8 vCPU, 16 GB RAM, dedicated CPU where offered
Block Storage 100 GB SSD
Object Storage 1 TB stored
Managed PostgreSQL 2 vCPU, 4 GB RAM, 50 GB storage
Managed Redis®* 1 vCPU, 1 GB RAM
Load Balancer TCP, 1 node
Serverless Functions 25M invocations, 512 MB, 150 ms
Free egress allowance What the provider includes each month, if anything
Egress beyond the allowance 1 TB out per month

Not every provider offers these exact configurations. When they don't, we pick the closest available option. For example, if a provider offers 14 GB and 24 GB of RAM but not 16 GB, we pick the 14 GB option.

How we place VPS plans into Small, Medium, and Large

Specs come first. For each size, we list the provider's plan that's closest to the reference configuration. Price doesn't move a plan out of its size: a 2 vCPU plan belongs in Small even if it's expensive, and the table will show that.

Some providers don't offer anything near the Small reference. Their entry-level plan might already start at 4 vCPU. We may still list that plan in the Small column when both of these are true:

  1. Its specs aren't far above the category (roughly up to 4 vCPU and 8 GB RAM).
  2. Its price is in line with the other Small entries.

We do this because the tables are meant to show how much you get at comparable price points. If a provider's cheapest plan gives you 4 vCPU for the price of a 1 vCPU plan elsewhere, that's a comparison worth showing.

Each plan appears in one size only. When a provider's entry-level plan moves into Small this way, their next tiers map to Medium and Large. This means the same column can show different specs across providers, depending on where a plan sits within that provider's own lineup.

If a provider has no plan reasonably close to a reference and no plan that qualifies for the smaller category by price, we leave that entry out rather than force a bad match.

Reference specifications

The exact configuration we price for each service. Where a provider doesn't offer one, we pick the closest option rather than force a bad match.

VPS Instances and Managed Containers

  • x86_64 architecture, if available
  • Linux (Ubuntu)
  • No minimum ephemeral storage

Block Storage

  • At least 100 GB of SSD storage
  • At least 3,000 IOPS, if available
  • At least 100 MB/s throughput, if available
  • No snapshots

Object Storage

  • Provider's standard storage class (no infrequent access or reduced redundancy)
  • 1 TB of data stored
  • 100 GB of data transferred in
  • 250 GB of data transferred out
  • 1,000,000 GET requests
  • 100,000 PUT or POST requests
  • 1,000 lifecycle transition requests

Managed PostgreSQL

  • 1 node, 2 vCPU, 4 GB RAM
  • At least 50 GB storage
  • No replication
  • No minimum backup retention
  • 10 GB read per month
  • 5 GB written per month

Managed Redis®*

  • 1 node, 1 vCPU, 1 GB RAM
  • 10 requests per second
  • 1 KB average request size

Load Balancers

  • TCP protocol
  • At least 1 node, 1 vCPU, 1 GB RAM
  • At least 200 Mbps throughput
  • 1,000 new connections per second
  • 1 second average connection duration
  • 1 TB data processed per month
  • 3 targets, 3 rules

Serverless Executions

  • 25M invocations
  • 150 ms of execution time
  • 512 MB of memory

Egress / Outbound Data Transfer

  • 1 TB of data transferred out per month
  • The destination region is the same as the source region (either US East or EU Central, see assumptions above)
  • Using the provider's default network class
  • No free tier allowances or discounts

GPU pricing

We track 4635 GPU configurations from 79 providers, covering 107 GPU models.

GPU Instances

  • One listing is one instance: a GPU model, how many of them, and the CPU, RAM and disk that come with it
  • The hourly price is the provider's published rate for the whole instance
  • Monthly and per-second rates are normalized to hourly. We use the provider's own stated hours per month where they publish one, and the 720-hour month above otherwise
  • Not every hourly rate can be rented by the hour. Where a provider sells only a monthly or longer commitment, the row's billing type says so
  • We also show a per GPU per hour rate: the instance price divided by the GPU count, so an 8x H100 at $24.00 /hr is $3.00 /GPU/hr
  • Price is one of several ranking factors, not the ranking. Each pricing panel explains its own order

GPU Billing Types

  • On-demand: pay as you go, no commitment
  • Reserved: a lower rate in exchange for a minimum term. Where a provider sells more than one committed rate we quote the plainer one, a standard reserved instance rather than a convertible one at AWS, a reserved instance rather than a savings plan at Azure. We quote the no-upfront rate
  • Spot: a lower rate in exchange for the provider being able to reclaim the instance at any time
  • Custom: quote-only, with no published rate to compare

What a GPU Rate Covers

  • The instance itself: its GPUs, vCPU and RAM, and any storage that comes with it
  • Nothing optional: block storage, boot disks, egress, public IPs, support and software licences are extra
  • No free allowances subtracted
  • Providers split the bill differently and we follow what each one publishes. Google Cloud prices the GPU, the vCPU and RAM, and the attached local SSD separately, so an a3-highgpu-8g is all four added up
  • We price the plain Linux machine: shared tenancy, no bundled database or Windows licence, no dedicated-host or capacity-reservation variant

GPU Availability

  • In stock: available at the time of verification
  • Waitlist: expected to be available soon
  • Out of stock: unavailable at the time of verification
  • Not reported: the provider publishes no status

GPU Specs

  • Specs are resolved per form factor, not per model
  • An H100 PCIe has roughly 2 TB/s of memory bandwidth and no NVLink; the SXM part has 3.35 TB/s and 900 GB/s of NVLink

LLM pricing

We track 510 LLM listings from 26 providers, covering 185 models. The full table is the LLM API pricing page; the LLM cost calculator prices a workload across all of them.

LLM Rates

  • One listing is one model at one provider, with an input rate and an output rate, both in USD per 1M tokens
  • We list the base-tier, on-demand rate: no cached-input discounts, batch pricing, long-context surcharges, or prepaid and member tiers
  • Where a provider publishes in another currency we convert at the same ECB reference rate as everything else on the site
  • Where a model has several providers, the price the tables quote is the cheapest provider's pair: input and output from that one listing

Open-Weight Models

  • A model whose creator publishes the weights and lists no first-party hosted rate is listed under the creator with no price, and marked "Open weights"
  • Third-party providers serving it are listed with their own rates

Cost Estimates

  • A cost is input tokens times the input rate plus output tokens times the output rate, and nothing else
  • The worked example on every model page is 10M input and 2M output tokens, a 5:1 ratio typical of chat and retrieval work
  • Words and pages convert at 1.33 tokens per English word and 300 words per page, OpenAI's rule of thumb; Claude 4.7 and later tokenize about 30% denser, and text in languages other than English can take 1.5x to 3x the tokens
  • Reasoning tokens, tool calls and images are billed by the provider on top and are not modelled

LLM Specs

  • Context window is the maximum input tokens per request, max output the maximum per response, both as the creator publishes them
  • Knowledge cutoff, modalities and capabilities are as published by the creator

LLM self-hosting and model fit

We use the formula below to estimate which models fit a given GPU.

Memory

  • Estimated VRAM is weights plus KV cache plus 2 GB of runtime overhead
  • Bytes per parameter are effective rates that account for quantization metadata and packing overhead, not the nominal width: 0.55 at 4-bit, 1.0 at 8-bit, 2.0 at 16-bit
  • KV cache per token is computed from each model's own architecture: its layer count, KV heads and head size. Where we have no per-model figure we use 250 KB per token, a conservative fallback
  • We assume an FP8 KV cache where the GPU supports it, FP16/BF16 on the rest, and FP16/BF16 at 16-bit weights on any GPU
  • Estimates assume one request at 32K tokens of context. Concurrent requests multiply the KV term, so eight requests at 32K need roughly the cache of one at 256K
  • A model fits when its estimated memory is within 90% of the GPU's advertised VRAM, leaving the rest as headroom for runtime allocations

Which rental fits

  • Rental prices use the median on-demand rate for the GPU, over the 720-hour month above. They are estimates, not quotes: what a provider charges depends on the node it sells, the term you commit to, and what it bills separately
  • GPU counts are rounded up to the node sizes providers typically sell: 1, 2, 4, 8 GPUs
  • We price nodes up to 8 GPUs. A model that needs more shows the GPU count it would take and no price, because past one node the cost is dominated by interconnect, which we do not price
  • For widest compatibility, the self-hosting table looks for matches on Nvidia GPUs only
  • These are memory-fit estimates. We do not guarantee runtime support, throughput, latency, or that a build at the displayed precision exists
  • Break-even is the API spend that equals the cheapest 4-bit rental at one request, at a 5:1 input-to-output ratio.