New B200 spot capacity is live in US East from $1.69 per GPU-hour. See availability
Major outage

API returning errors for all requests

The API was unavailable for 32 minutes. Running instances, the console and billing were not affected.

Reference 95843ba77d4c · Updated just now

Status
ResolvedResolved — last update 2 Jan 2026, 01:03 UTC
Started
2 Jan, 00:31Friday 2 January 2026 UTC
Duration
32 minEnded 2 Jan 2026, 01:03 UTC
Components
1API
Timeline

What we posted, as we posted it.

Resolved4 updates · UTC
  1. ResolvedThis incident has been resolved. A post-incident review will be published here.2 Jan, 01:03
  2. MonitoringThe API has been serving normally since 00:52 UTC. We are monitoring and reviewing the change that caused this.2 Jan, 00:52
  3. IdentifiedA load balancer configuration push removed every API backend from the healthy pool. The previous configuration has been restored.2 Jan, 00:43
  4. InvestigatingThe API is returning 502 errors for most requests. Running instances are not affected. We are investigating.2 Jan, 00:31

Post-incident review

published 5 Jan 2026, 10:03 UTC
Impact
The API was unavailable for 32 minutes. Running instances, the console and billing were not affected.
Root cause
A configuration template rendered an empty backend list when a region flag was misspelled; validation only checked the syntax, not the content.
What we changed
Load balancer pushes are now rejected when the healthy backend count would drop to zero, and they roll out region by region.
Credits
Accounts with affected instances were credited for the interrupted minutes under the SLA, without having to ask.

Affected components

Details

Type
Major outage
Started
2 Jan 2026, 00:31 UTC
Ended
2 Jan 2026, 01:03 UTC
Impact
The API was unavailable for 32 minutes. Running instances, the console and billing were not affected.
Reference
95843ba77d4c

Follow the Atom feed