Gmail went down on Tuesday, taking Twitter search down with it and prompting questions about the reliability of cloud-based services. The 100-minute, widespread outage left Gmail users in the communications dark.
Google apologized and offered a prompt explanation that sounds similar to the culprit for past Gmail outages. Google took some of its Gmail servers offline to perform routine upgrades. This isn’t unusual, as Google typically reroutes the traffic to other servers.
The problem arose when Google slightly underestimated the load that some of its recent changes put on its request routers. Ironically, those changes were intended to make Gmail more reliable. The request routers were overloaded and caused a ripple effect through other request servers that brought the lot to a screeching halt.
The Overarching Impact
Gmail was back online within two hours, but could the continued outages cause credibility problems for Google and its cloud-based apps? The answer lies in the frequency of the outages, according to Greg Sterling, principal analyst at Sterling Market Intelligence.
This isn’t the first Gmail outage and it probably won’t be the last, but since this one involved the popular micro-blogging service Twitter, it received more attention than past overloads. If Gmail continues to experience these types of issues, Sterling said, it could spell trouble for its web-based applications strategy.
“Although people are pretty invested in Gmail — it’s a very successful product that people rely on — the outages have a negative halo effect on Google’s apps cloud strategy,” Sterling said. “Google is quite aware of the importance of keeping Gmail up and running because it implicates other areas where they are trying to get people to rely on their apps.”
Google Learns Its Lesson
Well aware indeed. After apologizing and explaining the root cause of the Gmail outage, Ben Treynor, Google’s vice president of engineering and the…