Skip to main content

replicate · 34d ago

Delayed scaling due to node failure

Resolved

Timeline

Started: Aug 10, 2026, 9:05 PM UTC

Last update: Aug 10, 2026, 9:45 PM UTC

Resolved: Aug 10, 2026, 9:45 PM UTC

Updates

  • resolved

    All affected models have been scaling correctly for more than 30 minutes at this point, and we see no residual prediction queues. Thank you for your patience!

    Aug 10, 2026, 9:45 PM UTC

  • monitoring

    Scaling decisions were delayed by nearly 1 hour after the controller responsible for emitting queue metrics failed to schedule on a soft-failed node. We have since cordoned and drained the node, and the controller is emitting queue metrics again.

    Aug 10, 2026, 9:05 PM UTC

← All outage history