Every number on this site is derived from a price a cloud provider published. This page states exactly which prices are included, how they are turned into a per-GPU-hour figure, and what the resulting benchmark can and cannot be used for. If a step here looks wrong, the number it produces is wrong, and that is the point of writing it down.
The GPUQuant Reference Price is a published list-price indicator. It is not a transaction-weighted market price, it is not a settlement index, and it does not reflect negotiated or committed-use discounts.
GPUQuant tracks first-party published prices only, from Amazon Web Services, Microsoft Azure and Google Cloud. Aggregated third-party pricing services and GPU marketplaces are deliberately out of scope: their numbers cannot be verified against a primary source.
The measured price is, without exception:
Four GPUs are published: NVIDIA H100 80GB SXM, NVIDIA H200 141GB SXM, NVIDIA B200 180GB, NVIDIA A100 80GB SXM. Each has its own page with the full series and the offerings behind it.
The same GPU name can cover materially different products, and blending them would be the fastest way to produce a wrong number. GPUQuant defines one canonical variant per model and publishes only that.
Across all four models the canonical unit is the 8-GPU SXM or HGX node. Holding node size constant is what makes the three providers comparable: a single-GPU shape carries a different share of host CPU and memory per GPU, so mixing shapes would compare different machines wearing the same GPU name.
Included: AWS p5, Azure ND H100 v5 and Google Cloud a3-highgpu-8g. The same silicon in the same NVLink-connected form factor, which is what makes a cross-provider median meaningful.
Included: AWS p5en and p5e, Azure ND H200 v5 and Google Cloud a3-ultragpu-8g.
Included: AWS p6-b200 and Google Cloud a4-highgpu-8g.
Included: AWS p4de, Azure NDm A100 v4 and Google Cloud a2-ultragpu-8g.
Where a mapping is uncertain, the offering is excluded from every published aggregate and logged on the data coverage page. Nothing is guessed into a benchmark.
No provider publishes a price for one GPU for one hour. All three publish something else, and each needs its own arithmetic.
AWS prices the whole instance. GPUQuant divides the Linux on-demand hourly rate by the GPU count in the AWS price file, which matches the count in the AWS instance specification tables.
Azure prices the whole virtual machine and its size name does not encode a GPU count, so GPUQuant divides by a count read from the Microsoft size-series documentation.
Google prices accelerators, vCPU and memory as separate SKUs, so GPUQuant rebuilds the complete machine price from its components before dividing by GPU count. Anything less would compare a bare accelerator against two whole machines.
If an eight-GPU node costs $24 per hour:
$24.00 complete node price ÷ 8 GPUs = $3.00 per GPU-hour $3.00 × 730 hours = $2,190 per GPU-month
Google Cloud is the case that needs care. It bills the accelerator, the vCPU and the memory as three separate SKUs, so the accelerator price on its own is not the price of renting the machine. Using it directly would compare a bare GPU against two complete machines and understate Google Cloud by roughly a tenth. GPUQuant rebuilds the complete machine price from the published components first:
a3-highgpu-8g, one region 8 GPUs × H100 accelerator SKU + 208 vCPU × A3 Instance Core SKU + 1872 GB × A3 Instance Ram SKU = complete machine price per hour ÷ 8 GPUs = USD per GPU-hour
The A4 family, which carries B200, is the exception: Google publishes a bundled per-GPU slice price and no A4 Core or Ram SKU exists, so the slice price is already a complete per-GPU figure and is used as published. Every stored observation records which method produced it and keeps the component values, so any figure can be taken apart again.
Google Cloud figures exclude the local SSD attached to A2 and A3 machines, which is billed separately. AWS and Azure include local NVMe in the instance price. The gap is around one percent of the node price.
Region is a real pricing dimension, not noise: the same H100 costs materially different amounts in different regions. GPUQuant never publishes a bare number without its regional scope.
Balancing at the provider level is the load-bearing step. Google Cloud publishes H100 in far more regions than AWS does, so a flat median across all quotes would be a Google Cloud index wearing a global label. Each provider counts once.
AWS regional prices → median = A Azure regional prices → median = B GCP regional prices → median = C Reference = median(A, B, C) Regional range = min(all), max(all)
With two providers the median is their average; with one it is that provider's own median, and the label says so. The provider count is stored with every published point and shown in the tooltip.
A GPU-month is 730 hours of continuous access to one GPU. It is a presentation of the same number, not a separate measurement:
GPU-month = GPU-hour × 730 $3.20 per GPU-hour → $2,336 per GPU-month
730 hours is the convention used for compute contracts, including the planned rental-index futures, which is why GPUQuant uses it rather than an actual calendar month length.
The resolution is monthly, because that is the finest the sources support. Nothing on this site is daily or intraday, and no observation is invented to fill a gap.
A step line holds the last published price until the next publication, which is how a list price behaves. A tooltip marks a held month with a dot so a held price is never mistaken for a fresh one.
AWS publishes a dated version of the full price list roughly once a month, archived back to December 2015. Every GPUQuant point is one of those official versions.
Azure publishes current prices only, with no historical archive of any kind. The Azure series begins when GPUQuant took its first snapshot and cannot be backfilled.
Google Cloud answers historical queries one calendar month at a time, back to January 2017. Every GPUQuant point is the price Google reported as effective in that month.
Each record keeps its own type, and the three are never blurred: an official publication from an archived AWS price list, an effective interval Google reported for that month, or a GPUQuant snapshot of Azure's current-only feed.
Series therefore start at different dates and have different lengths. A B200 series is short because B200 reached the clouds recently. An Azure series is short because Azure publishes no archive. Neither is an error, and charts show them at their real length rather than stretching them.
A provider with no eligible offering is absent, not zero. It renders as “not currently offered” with the reason, never as $0.00, an empty line or a gap that implies a price fell.
The clearest current example is B200 on Azure. Azure publishes GB200 NVL, a rack-scale Grace-Blackwell system in which Blackwell GPUs are coupled to Grace CPUs over NVLink. That is a different machine from an eight-GPU HGX B200 node, so Azure contributes no B200 quote rather than a misleading one.
When the contributing provider set changes between two months, the reference price can move for a reason that has nothing to do with prices. Every stored point records its provider set so that case can always be identified.
A period change is only published when both endpoints are measuring the same thing. GPUQuant rebuilds both endpoints from the providers present in both months and states that basis next to the number.
If no provider spans both endpoints, or the series does not reach back far enough, the answer is N/A with an explanation rather than a number that would quietly measure a coverage change.
This is why a twelve-month change can read N/A on a GPU whose price is plainly visible on the chart. The chart is showing the headline reference; the change is refusing to compare a two-provider month against a three-provider month.
Ingestion runs monthly, matching the publication cadence of the sources. AWS and Google Cloud are re-read for the newest month; Azure is snapshotted, which is the only way its series can ever grow.
Last successful ingestion: AWS 2026-08-29 · AZURE 2026-08-29 · GCP 2026-08-29
Everything currently excluded, unresolved or unclassified is listed on the data coverage page, including the exact SKUs and the reason each one is out.