Uneven coverage is a property of the sources, not a defect to be smoothed over. This page states where every series starts, which provider publishes which GPU, and every offering GPUQuant deliberately leaves out.
Availability of the canonical eight-GPU node, with the number of regions in the newest month and the month each provider series begins.
| GPU | AWS | Azure | Google Cloud | Earliest usable history |
|---|---|---|---|---|
| H100 H100 80GB SXM5, eight-GPU node | Available 13 regions · from Jul 2023 | Available 25 regions · from Aug 2026 | Available 41 regions · from Feb 2024 | Jul 2023 |
| H200 H200 141GB SXM5, eight-GPU node | Available 11 regions · from Dec 2024 | Available 31 regions · from Aug 2026 | Available 6 regions · from Jan 2025 | Dec 2024 |
| B200 B200 180GB HGX, eight-GPU node | Available 4 regions · from Jun 2025 | Not offered Azure publishes GB200 NVL, a rack-scale Grace-Blackwell system, but no standalone eight-GPU B200 node. Blending the two would compare different machines. | Available 5 regions · from Jun 2026 | Jun 2025 |
| A100 80GB A100 80GB SXM4, eight-GPU node | Available 6 regions · from May 2022 | Available 32 regions · from Aug 2026 | Available 41 regions · from Oct 2022 | May 2022 |
Ingestion runs monthly, matching how often the sources publish. A source that misses that window is flagged here and on the market page rather than quietly going stale.
AWS publishes a dated version of the full price list roughly once a month, archived back to December 2015. Every GPUQuant point is one of those official versions.
Azure publishes current prices only, with no historical archive of any kind. The Azure series begins when GPUQuant took its first snapshot and cannot be backfilled.
Google Cloud answers historical queries one calendar month at a time, back to January 2017. Every GPUQuant point is the price Google reported as effective in that month.
Where a mapping cannot be confirmed against provider documentation, the offering is kept out of every published aggregate and recorded here. Nothing is guessed into a benchmark.
Grace-Blackwell rack-scale system. Microsoft publishes no accelerator count GPUQuant could verify, and it is not an eight-GPU B200 node in any case. Excluded rather than guessed.
A flex variant of the GB200 NVL system. Excluded for the same reason.
Previous generation, not tracked.
A100 40GB. Excluded: half the memory of the tracked variant.
specification ↗Blackwell Ultra, 2144 GiB accelerator memory across eight GPUs. Recorded, not part of the v1 model set.
specification ↗Grace-Blackwell rack-scale system, not an eight-GPU B200 node.
specification ↗Previous generation, not tracked.
Previous generation, not tracked.
L40S. Only AWS publishes it, so it cannot support a cross-provider benchmark.
Graphics family, not tracked.
Graphics family, not tracked.
Inference family, not tracked.
Fractional GPU shapes, not tracked.
Inference family, not tracked.
Fractional GPU shapes, not tracked.
Not tracked.
Not tracked.
Not tracked.
Not tracked.
Not tracked.
Not tracked.
AMD, not NVIDIA. The AWS GPU column is a count, not a vendor selector.
AWS silicon, not NVIDIA.
AWS silicon, not NVIDIA.
AWS silicon, not NVIDIA.
AWS silicon, not NVIDIA.
AWS silicon, not NVIDIA.
Intel Habana silicon, not NVIDIA.
Qualcomm silicon, not NVIDIA.
FPGA, not a GPU.
FPGA, not a GPU.
Video transcoding accelerator, not a GPU.
PCIe A100, excluded.
specification ↗ND A100 v4 is the 40GB part. Excluded from the 80GB benchmark.
specification ↗NC A100 v4 is the PCIe part with no SXM NVLink fabric. Excluded as a different variant.
specification ↗NCads H100 v5 is H100 NVL with 94GB. Different capacity, excluded.
specification ↗NCCads H100 v5, confidential computing, one H100 NVL 94GB. Excluded.
specification ↗PCIe A100, excluded.
specification ↗H100 NVL 94GB, excluded.
specification ↗Grace-Blackwell rack-scale system. Microsoft publishes no accelerator count GPUQuant could verify, and it is not an eight-GPU B200 node in any case. Excluded rather than guessed.
A flex variant of the GB200 NVL system. Excluded for the same reason.
Not tracked.
specification ↗Not tracked.
specification ↗Not tracked.
specification ↗Not tracked.
specification ↗A separately priced H100 tier. Including it would give Google Cloud two H100 quotes in one region.
The a3-megagpu networking tier, priced above standard H100.
Half the memory of the tracked variant.
Workstation-class part, not tracked.
Outside the v1 model set.
A fraction of a GPU is not a whole-GPU rental.
The Azure Retail Prices API returns current prices only. There is no version index and no historical endpoint, so the Azure series begins with the first GPUQuant snapshot and can never be backfilled. Its effectiveStartDate field says when the current price took effect, which is metadata about a standing price, not a time series.
None of the three providers returns the NVIDIA model in its pricing data. Every mapping is made by hand against vendor documentation and has to be rechecked when a new instance family launches.
Dynamic Workload Scheduler, Calendar mode and Reserved rows all carry usageType OnDemand at roughly half list price. Filtering on usage type alone would understate Google Cloud by around half, so those rows are excluded explicitly.
g4ad is AMD, inf2 is Inferentia and trn1 is Trainium, all with a non-zero GPU count. An explicit NVIDIA allowlist is required, and anything outside it is logged rather than assumed.
The column set grows from 61 fields in 2015 to 93 today. Filters are applied only where the column exists, because requiring a modern column against an old file matches nothing and looks exactly like a source with no data.
A GPU offered in six regions is measured on six regions. That makes each series correct for its own footprint but means a level comparison between two GPUs partly reflects where each is sold, not only what it costs.
The full calculation is on the methodology page, with worked examples in the calculations documentation.