RTX 4080 SUPER from $0.06 per GPU-hour
RTX 4080 SUPER is a consumer board rented by the hour with 16 GB of GDDR6X, built on the Ada Lovelace architecture in 2024. It holds models up to about 6B at FP16, and 13B-class models quantised.
Billed per started minute. 64 GB of GPU memory on a 4× node. Persistent storage $0.08 per GB-month. No egress fees.
- GPU memory
- 16GBGDDR6X
- Node sizes
- 1× · 2× · 4×Same price per GPU at every size
- Reclaim rate
- 10–15%Spot instances reclaimed over 30 days
What you get with each RTX 4080 SUPER node.
| Node | GPU memory | vCPU | RAM | Local NVMe | Spot | On-demand | Action |
|---|---|---|---|---|---|---|---|
| 1× RTX 4080 SUPER | 16 GB | 6 | 32 GB | 500 GB | $0.06/h | $0.09/h | Configure |
| 2× RTX 4080 SUPER | 32 GB | 12 | 64 GB | 1 TB | $0.12/h | $0.18/h | Configure |
| 4× RTX 4080 SUPER | 64 GB | 24 | 128 GB | 2 TB | $0.24/h | $0.36/h | Configure |
GPU-to-GPU link: PCIe Gen4 x16. Local NVMe is scratch space wiped when the instance ends; keep anything you need on a persistent disk.
64 RTX 4080 SUPER GPUs free right now.
Straight from the capacity pool the console books against. Prices are identical in every region — pick the one closest to your data.
A region with zero free GPUs still accepts reserved capacity requests — we hold hardware for a term rather than sell what is already taken.
What fits in 16 GB — and what it costs per hour.
Weights need about 2 GB per billion parameters at FP16 / BF16, 1 GB per billion parameters at FP8, 0.5 GB per billion parameters at INT4. We keep 20% of the memory free for activations, the KV cache and the CUDA context, then take the smallest node that still fits.
| Open model | Parameters | FP16 / BF16node · spot price | FP8node · spot price | INT4node · spot price |
|---|---|---|---|---|
| Mistral 7BAssistants, RAG, classification | 7.2B | 2× $0.12/h | 1× $0.06/h | 1× $0.06/h |
| Llama 3.1 8BAssistants, agents, fine-tuning | 8B | 2× $0.12/h | 1× $0.06/h | 1× $0.06/h |
| Gemma 2 27BHigher-quality assistants | 27B | over 4× | 4× $0.24/h | 2× $0.12/h |
| Qwen2.5 32BCode and reasoning | 32B | over 4× | 4× $0.24/h | 2× $0.12/h |
| Mixtral 8x7BMixture of experts, high throughput | 46.7B | over 4× | 4× $0.24/h | 2× $0.12/h |
| Llama 3.3 70BThe common production baseline | 70B | over 4× | over 4× | 4× $0.24/h |
| Qwen2.5 72BMultilingual, long context | 72B | over 4× | over 4× | 4× $0.24/h |
| Mixtral 8x22BMixture of experts, large capacity | 141B | over 4× | over 4× | over 4× |
| Llama 3.1 405BLargest widely used open model | 405B | over 4× | over 4× | over 4× |
Real usage depends on context length, batch size and the serving engine: a long context can add tens of gigabytes of KV cache. Treat the table as the floor, not the ceiling — and if a job is close to the limit, take the next node size or quantise one step further.
What a dollar buys on an RTX 4080 SUPER.
Ranked 10 of 14 consumer models in this catalogue on cost per gigabyte of GPU memory.
GPU memory you get for $1.00 of spot time, per hour.
Spot price divided by published FP32 throughput.
One GPU on spot, billed per minute, storage excluded.
Against the rest of the market
1 providers · September 2026- Cheapest on-demand seen
- $0.16Our on-demand $0.09 is 44% below it
- 24 hours on one GPU
- $1.44 on spot$2.16 on-demand · $3.84 at the cheapest rate found elsewhere
Index built from published prices for the same GPU across 137 providers (public price index (getdeploying.com)), September 2026. How the guarantee is enforced.
What people run on an RTX 4080 SUPER.
Quantised inference
Small and medium models in 4-bit or 8-bit form, with the whole model resident in memory.
Computer vision and batch jobs
Detection, classification, embeddings and transcoding at a very low hourly rate.
Anything that can checkpoint
On the spot tier this model is 33% below its own on-demand price, with a 2-minute notice before a reclaim.
NVIDIA RTX 4080 SUPER, as published by NVIDIA.
Taken from the vendor datasheet. Where a figure is not published, the row is absent rather than estimated.
- Vendor
- NVIDIA
- Architecture
- Ada Lovelace (2024)
- Segment
- Consumer
- Memory
- 16 GB GDDR6X
- FP32
- 52 TFLOPS
- CUDA cores
- 10,240
- Board power
- 320 W
- Form factor
- PCIe
- Interconnect
- PCIe Gen4 x16
- Multi-instance (MIG)
- Not available
- 8-bit float (FP8)
- Supported by the architecture
Source: www.nvidia.com · not confirmed on a vendor document
Software that runs on it
8 templatesPre-built environments, pulled on the node before you land on it. Or bring any OCI image from a public or private registry.
On the spot tier
10–15% reclaimed · 30 d- 2 minutes of notice on the metadata endpoint, a webhook and the console
- The instance is stopped, not deleted — disk and IP stay attached
- Auto-relaunch on the next free RTX 4080 SUPER, or switch the same disk to on-demand
- You save $0.72 a day per GPU against on-demand
If the RTX 4080 SUPER is not the right fit.
Same memory or more, 17% cheaper on spot
RTX 4000 Ada 20 GB GDDR6 · Ada Lovelace $0.06/h spotNext step up in memory: 20 GB instead of 16 GB
RTX 4090 24 GB GDDR6X · Ada Lovelace $0.06/h spotSame Ada Lovelace generation
RTX 4080 16 GB GDDR6X · Ada Lovelace $0.06/h spotSame Ada Lovelace generation
Or put two of them side by side: RTX 4080 SUPER vs RTX 3090 · RTX 4080 SUPER vs RTX 4000 Ada · RTX 4080 SUPER vs RTX 4090
RTX 4080 SUPER — the questions that come up
How much does an RTX 4080 SUPER cost per hour?
Spot is $0.06 per GPU-hour and on-demand is $0.09, both billed per minute — every started minute costs the hourly price divided by 60, so $0.0010 on spot. Reserved capacity is $0.06 per GPU-hour on a monthly commitment. Storage and egress are not included in that rate: persistent disks are $0.08 per GB-month and there are no egress fees.
How many GPUs can I put in one instance?
Node sizes are 1×, 2×, 4× — up to 4 RTX 4080 SUPER GPUs in a single virtual machine, with CPU, RAM and local NVMe scaled with the GPU count. The price per GPU is identical at every node size.
Can it run a 70B model?
Not on a single node of this model: at FP8 a 70B model needs about 70 GB of memory for the weights alone, more than 4× 16 GB leaves free. Quantised to 4 bits it needs about 35 GB — check the memory table above — or pick a model with more memory per GPU.
Which regions have it?
Available now in US West (Oregon), EU Central (Frankfurt). Prices are identical in every region; pick the one closest to your data.
How often is a spot RTX 4080 SUPER reclaimed?
Over the trailing 30 days, 10–15% of spot instances on this model were reclaimed. You get a 2-minute notice on the metadata endpoint, through a webhook and in the console; the instance is then stopped, never deleted, and the persistent disk stays attached.
Anything else about this model? Ask an engineer — the same people run the nodes.
An RTX 4080 SUPER for $0.06 an hour, running in under 60 seconds.
Pay as you go — no contracts, no minimum commitment. Add credit, launch, stop whenever you want.
Billed per minute from the moment the instance is reachable. Minimum credit $40, no subscription. 64 RTX 4080 SUPER GPUs available right now.