New B200 spot capacity is live in US East from $1.69 per GPU-hour. See availability
Use cases

Start from the job, not from the GPU.

Four workloads, each with the memory arithmetic, the cheapest model in the catalogue that fits, what an hour costs, and whether it belongs on interruptible capacity.

The common thread

Three questions decide every one of them.

Whatever the workload, the answer comes from the same three checks — which is why the guides all look the same underneath.

  1. Does it fit in memory?

    Weights need roughly 2 GB per billion parameters at FP16, 1 GB at FP8, 0.5 GB in 4-bit, plus about 20% for activations and cache. Every model page does the arithmetic for the exact card.

  2. Is it bound by bandwidth or by compute?

    Token generation reads the whole model from memory for every token, so bandwidth sets the ceiling. Training and image generation lean harder on compute.

  3. Can it survive a 2-minute warning?

    If yes, spot cuts the bill by 30–53%. If not, on-demand costs more and never stops. How spot works.

Get started

Your first H100 for $0.55 an hour.

Pay as you go — no contracts, no minimum commitment. Add credit, launch, stop whenever you want.

Billed per minute from the moment the instance is reachable. Minimum credit $40, no subscription.