Gemini (Google) · 2d ago
We are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region.
ResolvedTimeline
Started: Aug 20, 2026, 3:40 PM UTC
Last update: Aug 27, 2026, 9:45 PM UTC
Resolved: Aug 20, 2026, 7:20 PM UTC
Affected: Contact Center AI Platform
Updates
unknown
# Incident Report ## Summary On Thursday, 20 August 2026, from 08:00 to 10:22 US/Pacific (15:00 to 17:22 UTC), multiple Google Cloud services in the us-west1 region experienced elevated latency, provisioning failures, increased error rates, and widespread service degradations for a total duration of 2 hours and 22 minutes. The incident impacted a wide range of core services — including Persistent Disk, Google Kubernetes Engine, Google Compute Engine, Cloud Run, Google Cloud Bigtable, Cloud Storage, and Identity and Access Management — across both control plane operations and data plane requests. We sincerely apologize for the disruption this incident caused to your business operations and critical workloads. We recognize the vital role Google Cloud plays in supporting your organization, and we deeply regret the impact on your business operations and critical workloads. Engineering and infrastructure teams are actively implementing measures to address the root causes and strengthen network resiliency to prevent recurrences in the future. Specifically, teams are monitoring recovery progress, refining safety checks for planned optical maintenance, and optimizing regional traffic routing mechanisms to safeguard against unexpected capacity constraints. ## Root Cause The disruption originated during scheduled fiber optic maintenance, which unexpectedly compromised network capacity between data centers within the us-west1 region. Automated rerouting mechanisms failed to properly redistribute traffic to alternate capacity, resulting in network congestion as volumes exceeded available bandwidth in the affected area. This underlying network degradation subsequently impacted higher-level service components through severe packet loss, request throttling, and increased latency across inter-campus dependencies. Core infrastructure services, including Spanner Paxos consensus and the Unified Metadata Server (UMS), experienced significant latency spikes, which cascaded into timeouts and elevated error rates for downstream dependent products such as Cloud Storage, Cloud IAM, Persistent Disk, and Google Kubernetes Engine. Consequently, both control plane operations and data plane requests failed to execute successfully across multiple Google Cloud services in us-west1 throughout the duration of the incident. ## Remediation and Prevention Internal monitoring systems initially detected widespread service anomalies and alerted Google engineers, who promptly confirmed that multiple core services operating across the us-west1 region were severely impacted. In response to the immediate operational risks, engineering teams promptly executed emergency traffic draining protocols to reroute active service workloads away from the compromised inter-campus network infrastructure and minimize further customer disruption. Once teams restored inter-campus fiber network capacity, engineers systematically reintroduced production traffic back to us-west1 through controlled validation stages, ultimately confirming full service normalization across all impacted platforms. Longer term engineering remediations and architectural enhancements are actively being finalized and assigned to owning teams. Key focus areas include reforming scheduled maintenance protocols, establishing stricter safety checks and circuit redundancy requirements prior to routine maintenance, and refining automated regional traffic failover mechanisms to automatically handle sudden capacity degradation without incurring severe network congestion. ## Detailed Description of Impact On Thursday, August 20, between 08:00 and 10:22 US/Pacific, Google Cloud customers in the us-west1 region encountered elevated latency, provisioning failures, and increased error rates. * **Impact / Error Rates:** The reduced capacity caused request throttling, latency spikes, and increased retry volume. * **Scope Exclusion:** Services and workloads operating in regions other than us-west1 remained fully operational and unaffected. **Affected Services and Features** The following services experienced elevated latencies and/or increased error rates, across their respective data planes and control planes: * AlloyDB * Apache Kafka * Apigee Edge Public Cloud * Apigee X * Artifact Registry * BigQuery & BigQuery Data Transfer Service * Cloud Build * Cloud Data Fusion * Cloud Dataflow * Cloud Filestore * Cloud Key Management Service (KMS) * Cloud Monitoring * Cloud Run * Cloud SQL * Contact Center AI Platform * Dataproc Metastore * Google App Engine * Google Cloud Bigtable * Google Cloud Pub/Sub * Google Cloud Storage (GCS) * Google Compute Engine (GCE) * Google Kubernetes Engine (GKE) * Identity and Access Management (IAM) * Managed Airflow (Cloud Composer) * Managed Service for Apache Spark (Dataproc) * Persistent Disk The regional incident in us-west1 followed a structured three-phase recovery dictated by platform dependency layers: While core platform connectivity and live request serving were restored by 10:22 US/Pacific, certain services took additional time to fully recover due to asynchronous backlog processing and localized control-plane state reconciliation for a very small set of customers. For event-driven and pipeline services, inbound error rates dropped to 0 immediately, but some operations experienced elevated latency while workers processed through backlogs accumulated during the outage, clearing later for Cloud Pub/Sub, Cloud Build & Deploy, Cloud Dataflow, and Cloud Storage lifecycle deletions. Concurrently, while primary read/write traffic was healthy across the region, specific long-running lifecycle workflows required extra time to clear locks and reconcile distributed state machines. Compute Engine VM provisioning in zone us-west1-c and Cloud Filestore control plane instance allocation locks and resource validation checks took longer to normalize across regional storage backends.
Aug 27, 2026, 9:45 PM UTC
unknown
# Preliminary Incident Report We sincerely apologize for the disruption this incident caused to your business operations. Recognizing your reliance on Google Cloud, we express our sincere regrets for any operational impact experienced. Our engineering teams are actively addressing the underlying root cause to prevent future recurrences. Please note that the information provided herein reflects our current understanding as of the time of publication and remains subject to revision as the investigation progresses. A comprehensive Incident Report detailing preventive measures will be issued upon conclusion of our inquiry. If you have experienced impact outside of what is listed below, please reach out to Google Cloud Support using https://cloud.google.com/support. ## Date/Time of the Issue (All time US/Pacific) * Incident Start: 20 August 2026 08:00 * Incident End: 20 August 2026 10:22 * Duration: 2 hours, 22 minutes ## Summary On Thursday, 20 August 2026, multiple Google Cloud services encountered elevated latency, provisioning failures, increased error rates, and service degradations lasting for a duration of 2 hours and 22 minutes. ## Preliminary Root Cause The disruption originated during scheduled fiber optic maintenance, which unexpectedly compromised network capacity between data centers within the us-west1 region. Automated rerouting mechanisms failed to properly redistribute traffic to alternate capacity, resulting in network congestion as volumes exceeded available bandwidth in the affected area. This underlying network degradation subsequently impacted higher-level service components through request throttling, increased latency, and cascading retries. Consequently, both control plane operations and data plane requests failed to execute successfully across several dependent Google Cloud services, producing the observed latency and elevated error rates. ## Remediation Internal monitoring alerted Google engineers to the incident, confirming that multiple services across us-west1 were affected. To mitigate the immediate operational impact, engineers initiated traffic draining protocols, re-routing service workloads away from the degraded network infrastructure. Following the restoration of inter-campus network capacity and the resolution of the core issue, engineers systematically restored traffic to the us-west1 region and confirmed service normalization. Long-term engineering solutions are currently being finalized and assigned to address the root cause and strengthen system resilience regarding fiber maintenance procedures. ## Description of Impact On Thursday, August 20, between 08:00 and 10:22 US/Pacific, Google Cloud customers in the us-west1 region encountered elevated latency, provisioning failures, and increased error rates. * Proportionate Impact / Error Rates: The precise scope of impact—including the proportion of affected projects and applications—and detailed error metrics will be published in the comprehensive Root Cause Analysis, as the reduced capacity caused request throttling, latency spikes, and increased retry volume. * Scope Exclusion: Services and workloads operating in regions other than us-west1 remained fully operational and unaffected. **Affected Services and Features** The following services experienced elevated latencies and/or increased error rates, across their respective data planes and control planes: * AlloyDB * Apache Kafka * Apigee Edge Public Cloud * Apigee X * Artifact Registry * BigQuery & BigQuery Data Transfer Service * Cloud Build * Cloud Data Fusion * Cloud Dataflow * Cloud Filestore * Cloud Key Management Service (KMS) * Cloud Monitoring * Cloud Run * Cloud SQL * Contact Center AI Platform * Dataproc Metastore * Google App Engine * Google Cloud Bigtable * Google Cloud Pub/Sub * Google Cloud Storage (GCS) * Google Compute Engine (GCE) * Google Kubernetes Engine (GKE) * Identity and Access Management (IAM) * Managed Airflow (Cloud Composer) *Managed Service for Apache Spark (Dataproc) Persistent Disk
Aug 25, 2026, 6:20 AM UTC
unknown
**Summary** The issue causing timeouts, degradations, errors, and latencies across multiple us-west1 products has been mitigated. **Description** We have mitigated the issue impacting multiple products in our us-west1 region as of Thursday, 2026-08-20 10:22 PDT. Our engineering teams have restored capacity on a planned optical maintenance that caused unexpected congestion in Dalles, Oregon metro /us-west1 region. Our systems stabilized and services have recovered once capacity was restored. Our teams are continuing to monitor for any residual impact. We will publish an analysis of this incident once we have completed our internal investigation. We thank you for your patience while we worked on resolving the issue. **Diagnosis / Customer Symptoms** Customers in us-west1 may have experienced timeouts, service degradations, errors, and elevated latencies across multiple products. **Workaround** This issue is now mitigated.
Aug 20, 2026, 7:37 PM UTC
unknown
**Summary** The issue causing timeouts, degradations, errors, and latencies across multiple us-west1 products has been mitigated and we are working to recover all products. **Description** Mitigation actions have been completed by our engineering teams and we are seeing recovery from multiple products. We are continuing to work to recover the remaining products. We will provide an update by Thursday, 2026-08-20 13:30 PDT with details. **Diagnosis / Customer Symptoms** Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. **Workaround** No workarounds needed at this time.
Aug 20, 2026, 7:21 PM UTC
unknown
**Summary** The issue causing timeouts, degradations, errors, and latencies across multiple us-west1 products has been mitigated and we are working to recover all products. **Description** Mitigation actions have been completed by our engineering teams and we are seeing recovery from multiple products. We are continuing to work to recover the remaining products. We will provide an update by Thursday, 2026-08-20 12:30 PDT with details. **Diagnosis / Customer Symptoms** Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. **Workaround** No workarounds needed at this time.
Aug 20, 2026, 6:53 PM UTC
unknown
**Summary** The issue causing timeouts, degradations, errors, and latencies across multiple us-west1 products has been mitigated and we are working to recover all products. **Description** Mitigation actions have been completed by our engineering teams and we are seeing recovery from multiple products. We are continuing to work to recover the remaining products. We will provide an update by Thursday, 2026-08-20 11:45 PDT with details. **Diagnosis / Customer Symptoms** Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. **Workaround** No workarounds needed at this time.
Aug 20, 2026, 6:03 PM UTC
unknown
**Summary** We are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region. **Description** Mitigation actions are implemented by our engineering teams, recovery trends have been observed across infrastructure layers. Active efforts remain underway to bring impacted cloud services back to full operation. We will provide an update by Thursday, 2026-08-20 11:00 PDT with details. **Diagnosis / Customer Symptoms** Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. **Workaround** We recommend customers to failover to other regions where feasible.
Aug 20, 2026, 5:32 PM UTC
unknown
**Summary** We are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region. **Description** Mitigation actions are implemented by our engineering teams, recovery trends have been observed across infrastructure layers. Active efforts remain underway to bring impacted cloud services back to full operation. We will provide an update by Thursday, 2026-08-20 10:45 PDT with details. **Diagnosis / Customer Symptoms** Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. **Workaround** We recommend customers to failover to other regions where feasible.
Aug 20, 2026, 5:13 PM UTC
unknown
**Summary** We are investigating an issue where customers may experience timeouts, service degradations, errors, and elevated latencies across multiple products in the us-west1 region. **Description** We are experiencing an issue with multiple products, beginning on Thursday, 2026-08-20 08:40 PDT. Our engineering team continues to investigate the issue. We will provide an update by Thursday, 2026-08-20 10:30 PDT with details. We apologize to all who are affected by the disruption. **Diagnosis / Customer Symptoms** Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. **Workaround** None at this time.
Aug 20, 2026, 4:44 PM UTC