3 min read

Google Cloud outage exposes hidden single-site risks

A 15-hour Google Cloud outage hit three services in europe-west4-a after a power and cooling failure, raising fresh questions about hyperscaler resilience.

Image: The Register

A 15-hour Google Cloud outage last week took down three services in europe-west4-a while the rest of the zone and region stayed online, underscoring how hard it can be for customers to judge the real resilience of hyperscaler infrastructure.

According to Google’s incident report, Google Cloud VMware Engine (GCVE), NetApp Volumes, and Bare Metal Solution (BMS) were affected by a cooling failure. Google said “The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure.”

That detail matters because it shows those three services depend on a discrete datacenter. Google also said “An electrical fault occurred on the utility grid upstream of the datacenter, disrupting the electrical distribution gear and cooling equipment.” The company has not yet explained how the upstream grid issue caused that disruption, but said it “proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment.”

Recommended reading

UK data centre boom hits a water reality check

The Register said it asked Google whether the site had generators or other backup power sources, and if so why workloads still had to be shut down. At the time of writing, Google had not responded. The company said its incident analysis is currently ongoing and promised a follow-up.

Hidden single-datacenter dependencies

The outage has drawn attention to a familiar cloud assumption: that deploying across multiple zones and regions is enough to guarantee resilience. Google, like other hyperscalers, structures its cloud into regions made up of multiple zones and typically advises customers to spread workloads across them. But this incident suggests some managed services can still be pinned to a single datacenter inside a zone.

Forrester principal analyst Biswajeet Mahapatra told The Register that transparency is the real issue.

“The real issue is transparency: customers are generally told to use multiple zones and regions for resilience but are rarely given visibility into whether a particular managed service has a single-datacenter dependency within a zone.” “As a result, many organizations assume the cloud abstraction provides more facility-level redundancy than may actually exist for specialized services.”

Biswajeet Mahapatra, Principal Analyst, Forrester

Mahapatra added that the architecture itself is not unusual.

“The underlying architecture is not necessarily unusual. AWS, Azure, and Google all operate services that rely on dedicated hardware, storage platforms, or tightly coupled infrastructure that may not be distributed across multiple facilities in the same way as core compute and storage services.”

Biswajeet Mahapatra, Principal Analyst, Forrester

What the outage reveals about region design

Adrian Wong, a Gartner Director Analyst, pointed The Register to Google Cloud’s 2023 outage in europe-west9-a, which Google attributed to a water leak that “originated in a non-Google portion of the facility.” Google uses Spanner to replicate data across zones, but Wong noted that setup did not hold once one building in the flooded zone became unavailable.

“It is very hard to figure out how an individual region is architected.” “Our customers are often surprised by that.”

Adrian Wong, Director Analyst, Gartner

Google’s incident report includes an apology, saying: “We know how much you rely on Google Cloud, and we regret the impact on your productivity,” and promising a final report with “preventative actions.” For customers, though, the tougher question may be whether those actions also address the hidden design choices that can quietly limit resilience.

Marcus Vance

Enterprise Editor

Marcus follows the money. He covers enterprise software, cloud architecture, and the tectonic shifts in Big Tech strategy. He translates dense earnings calls and complex M&A activity into actionable insights about where the industry is actually heading. If a tech giant makes a silent pivot, Marcus is usually the first to notice.

via The Register

// Keep reading