New B200 spot capacity is live in US East from $1.69 per GPU-hour. See availability
NVIDIA · Hopper · Data-center GPU

H200 NVL from $0.69 per GPU-hour

An H200 on a PCIe card with an NVLink bridge. Slightly lower clocks than the SXM board, the same 141 GB of memory, and a lower price — a good fit when you need the capacity more than the last few percent of throughput.

Specifications from the vendor datasheet · prices updated recently

Price per GPU-hour 16 available
SpotReclaimed with 2 min notice · 5–10% over 30 days $0.69$0.0115 / min
On-demandRuns until you stop it $0.99−30% on spot
ReservedFixed rate, 1 to 12 months $0.69monthly term

Billed per started minute. 1128 GB of GPU memory on a 8× node. Persistent storage $0.08 per GB-month. No egress fees.

GPU memory
141GBHBM3e · MIG capable
Bandwidth
4,800GB/sSets token throughput on inference
Node sizes
1× · 2× · 4× · 8×Same price per GPU at every size
Reclaim rate
5–10%Spot instances reclaimed over 30 days
Node sizes

What you get with each H200 NVL node.

All prices
Node GPU memory vCPU RAM Local NVMe Spot On-demand Action
1× H200 NVL 141 GB 26 240 GB 2 TB $0.69/h $0.99/h Configure
2× H200 NVL 282 GB 52 480 GB 4 TB $1.38/h $1.98/h Configure
4× H200 NVL 564 GB 104 960 GB 8 TB $2.76/h $3.96/h Configure
8× H200 NVL 1128 GB 208 1,920 GB 16 TB $5.52/h $7.92/h Configure

GPU-to-GPU link: NVLink bridge 900 GB/s per GPU (2- or 4-way); PCIe Gen5 128 GB/s. Local NVMe is scratch space wiped when the instance ends; keep anything you need on a persistent disk.

Availability

16 H200 NVL GPUs free right now.

Straight from the capacity pool the console books against. Prices are identical in every region — pick the one closest to your data.

US EastVirginia, US 16
Tier III100 Gbps per nodeus-east

A region with zero free GPUs still accepts reserved capacity requests — we hold hardware for a term rather than sell what is already taken.

Memory planning

What fits in 141 GB — and what it costs per hour.

Weights need about 2 GB per billion parameters at FP16 / BF16, 1 GB per billion parameters at FP8, 0.5 GB per billion parameters at INT4. We keep 20% of the memory free for activations, the KV cache and the CUDA context, then take the smallest node that still fits.

Open model Parameters FP16 / BF16node · spot price FP8node · spot price INT4node · spot price
Mistral 7BAssistants, RAG, classification 7.2B $0.69/h $0.69/h $0.69/h
Llama 3.1 8BAssistants, agents, fine-tuning 8B $0.69/h $0.69/h $0.69/h
Gemma 2 27BHigher-quality assistants 27B $0.69/h $0.69/h $0.69/h
Qwen2.5 32BCode and reasoning 32B $0.69/h $0.69/h $0.69/h
Mixtral 8x7BMixture of experts, high throughput 46.7B $0.69/h $0.69/h $0.69/h
Llama 3.3 70BThe common production baseline 70B $1.38/h $0.69/h $0.69/h
Qwen2.5 72BMultilingual, long context 72B $1.38/h $0.69/h $0.69/h
Mixtral 8x22BMixture of experts, large capacity 141B $2.76/h $1.38/h $0.69/h
Llama 3.1 405BLargest widely used open model 405B $5.52/h $2.76/h $1.38/h
This is an estimate for weights, not a benchmark.

Real usage depends on context length, batch size and the serving engine: a long context can add tens of gigabytes of KV cache. Treat the table as the floor, not the ceiling — and if a job is close to the limit, take the next node size or quantise one step further.

Value

What a dollar buys on an H200 NVL.

Ranked 14 of 21 data-center models in this catalogue on cost per gigabyte of GPU memory.

VRAM per dollar
204 GB

GPU memory you get for $1.00 of spot time, per hour.

Memory bandwidth per dollar
6,957 GB/s

Bandwidth decides token throughput far more often than raw FLOPS.

Cost per FP16 TFLOP-hour
$0.00083

Spot price divided by published FP16 tensor throughput.

Cost of a 24-hour run
$16.56

One GPU on spot, billed per minute, storage excluded.

Against the rest of the market

6 providers · September 2026
Median on-demand elsewhere
$3.83Our on-demand is 74% below it
Cheapest on-demand seen
$2.45Our on-demand $0.99 is 60% below it
Cheapest spot seen
$1.32Our spot $0.69 is 48% below it
24 hours on one GPU
$16.56 on spot$23.76 on-demand · $58.80 at the cheapest rate found elsewhere

Index built from published prices for the same GPU across 137 providers (public price index (getdeploying.com)), September 2026. How the guarantee is enforced.

Good for

What people run on an H200 NVL.

Training and full fine-tuning

Enough memory for optimiser states and activations on models a smaller card can only run in inference.

Serving large models

A 70B model at FP8 on a single GPU, with room for a long context window.

Multi-GPU jobs over NVLink

GPU-to-GPU traffic stays off the PCIe bus, which is what makes tensor and pipeline parallelism worth it.

Partitioned serving (MIG)

The GPU can be split into isolated instances, each with its own memory and compute slice.

Specifications

NVIDIA H200 NVL, as published by NVIDIA.

Taken from the vendor datasheet. Where a figure is not published, the row is absent rather than estimated.

Vendor
NVIDIA
Architecture
Hopper (2024)
Segment
Datacenter
Memory
141 GB HBM3e
Memory bandwidth
4,800 GB/s
FP16 / BF16 tensor
836 TFLOPS
FP8 tensor
1,671 TFLOPS
FP32
60 TFLOPS
Board power
600 W
Form factor
PCIe dual-slot
Interconnect
NVLink bridge 900 GB/s per GPU (2- or 4-way); PCIe Gen5 128 GB/s
Multi-instance (MIG)
Supported by the GPU
8-bit float (FP8)
Supported by the architecture

Source: www.nvidia.com · verified against the vendor document

Software that runs on it

8 templates

Pre-built environments, pulled on the node before you land on it. Or bring any OCI image from a public or private registry.

PyTorch 2.5 · CUDA 12.4CUDA 12.4 basevLLM inferenceJupyterLabComfyUIOllamaTensorFlow 2.17Ubuntu 22.04 (driver only)

All templates

On the spot tier

5–10% reclaimed · 30 d
  • 2 minutes of notice on the metadata endpoint, a webhook and the console
  • The instance is stopped, not deleted — disk and IP stay attached
  • Auto-relaunch on the next free H200 NVL, or switch the same disk to on-demand
  • You save $7.20 a day per GPU against on-demand

How interruptions work

FAQ

H200 NVL — the questions that come up

How much does an H200 NVL cost per hour?

Spot is $0.69 per GPU-hour and on-demand is $0.99, both billed per minute — every started minute costs the hourly price divided by 60, so $0.0115 on spot. Reserved capacity is $0.69 per GPU-hour on a monthly commitment. Storage and egress are not included in that rate: persistent disks are $0.08 per GB-month and there are no egress fees.

How many GPUs can I put in one instance?

Node sizes are 1×, 2×, 4×, 8× — up to 8 H200 NVL GPUs in a single virtual machine, with CPU, RAM and local NVMe scaled with the GPU count. The price per GPU is identical at every node size.

Can it run a 70B model?

Yes. At FP8, a 70B model needs about 70 GB for the weights, so it fits on 1× H200 NVL ($0.69 per hour on spot) with roughly 20% of the memory left for activations and the KV cache.

Which regions have it?

Available now in US East (Virginia). Prices are identical in every region; pick the one closest to your data.

How often is a spot H200 NVL reclaimed?

Over the trailing 30 days, 5–10% of spot instances on this model were reclaimed. You get a 2-minute notice on the metadata endpoint, through a webhook and in the console; the instance is then stopped, never deleted, and the persistent disk stays attached.

Anything else about this model? Ask an engineer — the same people run the nodes.

Get started

An H200 NVL for $0.69 an hour, running in under 60 seconds.

Pay as you go — no contracts, no minimum commitment. Add credit, launch, stop whenever you want.

Billed per minute from the moment the instance is reachable. Minimum credit $40, no subscription. 16 H200 NVL GPUs available right now.