New B200 spot capacity is live in US East from $1.69 per GPU-hour. See availability
Major outage

Console unavailable

The console was unavailable for 127 minutes.

Reference f39fa4b1213a · Updated just now

Status
ResolvedResolved — last update 26 Jan 2025, 20:02 UTC
Started
26 Jan, 17:55Sunday 26 January 2025 UTC
Duration
2 h 7 minEnded 26 Jan 2025, 20:02 UTC
Components
1Console
Timeline

What we posted, as we posted it.

Resolved5 updates · UTC
  1. ResolvedThis incident has been resolved.26 Jan, 20:02
  2. MonitoringThe console has been serving normally since 19:34 UTC. We are monitoring.26 Jan, 19:34
  3. IdentifiedThe rollback is complete. The failing deploy shipped a dependency that was incompatible with the production PHP runtime.26 Jan, 18:26
  4. UpdateWe have narrowed the cause down and are applying a fix. Next update within 30 minutes.26 Jan, 18:12
  5. Investigatingcloud.spotgpus.com is returning errors. Running instances and the API are not affected. We are rolling back the latest deploy.26 Jan, 17:55

Post-incident review

published 31 Jan 2025, 05:02 UTC
Impact
The console was unavailable for 127 minutes.
Root cause
The staging environment ran a newer runtime than production, so the incompatibility was invisible before the deploy.
What we changed
Staging and production now build from the same image digest.
Credits
Accounts with affected instances were credited for the interrupted minutes under the SLA, without having to ask.

Affected components

Details

Type
Major outage
Started
26 Jan 2025, 17:55 UTC
Ended
26 Jan 2025, 20:02 UTC
Impact
The console was unavailable for 127 minutes.
Reference
f39fa4b1213a

Follow the Atom feed