New B200 spot capacity is live in US East from $1.69 per GPU-hour. See availability
Degraded performance

Elevated API error rates

About 1.3% of API requests failed for 32 minutes.

Reference e85c34bbab1f · Updated just now

Status
ResolvedResolved — last update 6 Jun 2025, 13:05 UTC
Started
6 Jun, 12:33Friday 6 June 2025 UTC
Duration
32 minEnded 6 Jun 2025, 13:05 UTC
Components
1API
Timeline

What we posted, as we posted it.

Resolved4 updates · UTC
  1. ResolvedThis incident has been resolved. Requests that failed were not billed and can be retried safely.6 Jun, 13:05
  2. MonitoringError rates returned to normal at 12:56 UTC. We are monitoring while the node is rebuilt and put back behind the load balancer.6 Jun, 12:56
  3. IdentifiedThe errors come from one API node that kept serving after a routine deploy with an exhausted database connection pool. The node has been removed from rotation.6 Jun, 12:44
  4. InvestigatingWe are investigating elevated error rates on the API. About 1.3% of requests are receiving 5xx responses. The console is not affected.6 Jun, 12:33

Affected components

Details

Type
Degraded performance
Started
6 Jun 2025, 12:33 UTC
Ended
6 Jun 2025, 13:05 UTC
Impact
About 1.3% of API requests failed for 32 minutes.
Reference
e85c34bbab1f

Follow the Atom feed