Article 78H83 Your cloud survived everything except the real world

Your cloud survived everything except the real world

by
from www.theregister.com - Articles on (#78H83)
Story ImageYou don't know what you've got till it's gone. Great lyric, lousy data retention policy. Amazon Web Services said last week that war damage to its Middle East infrastructure had permanently destroyed resources and data hosted exclusively in its now rather badly named Bahrain Availability Zones. The damage overwhelmed the resilience built into the region. Customers without copies elsewhere no longer had their data. Sorry about that. This may have surprised anyone who mistook cloud redundancy for an intrinsic guarantee of safety. AWS is far from the only American operation to have suffered in the region: the US Navy has reportedly had its local maintenance and supply network badly mauled, with serious consequences for its operations. If systems designed to withstand war cannot cope with sustained physical attacks, civilian bit barns have little chance. The episode also underlines a familiar but easily neglected lesson: resilience within one cloud region is not the same thing as maintaining an independent backup elsewhere. Physical destruction is not the only threat. A major outage of the UK air traffic control system in September, which stranded hundreds of thousands of passengers and led to thousands of flight cancellations, was reportedly triggered by a military aircraft filing an incompatible flight plan. Presumably Flight Lieutenant Bobby Tables has been reprimanded. The apparent failure to validate the flight plan data was not the worst of it. NATS, which runs the UK's air traffic control system, reportedly told airports and airlines that no backup system was available because it was undergoing a "complete overhaul." Nor was there a backup of the live data, supposedly because of the "vast amounts" involved. If your system produces too much data to back up, you had better be running a particle collider or a giant telescope. Otherwise, you may be in the wrong business. More charitably, backup strategies are complicated, expensive, and difficult to test. They also suffer from the insurance problem: while nothing is going wrong, management sees only capital and operating expenditure with no obvious return. The AI infrastructure boom has made that problem considerably worse. A backup is a copy, and a copy needs storage. AI datacenter operators are swallowing much of the available capacity, with drives selling out faster than tickets for a Taylor Swift tour. Western Digital had allocated its entire 2026 hard-drive production run by mid-February. The only consolation for those responsible for keeping data safe is that they can say "Yeah? You buy it, then" to anyone who smugly invokes the 3-2-1 backup rule. Three copies on two different media with one kept off-prem? Lovely idea, if you don't have to provision it. Someone is provisioning it for all those giant datacenters that will run our lives, right? Right? Even outside the immediate reach of drones and missiles - a distinction that feels less reassuring with every passing month - infrastructure is operating in increasingly hostile conditions. Cables get cut, climate goes chaotic, criminals encrypt, commanders-in-chief go crazy. This makes planning for and implementing a sound data resilience strategy very hard, at exactly the same time as it becomes more important. Inter-cloud data duplication gets more complicated if digital sovereignty is a factor, especially if your sanctified region is in the firing line. Commerce has confronted a similar problem before, if we extend the backup-as-insurance metaphor. The development of marine insurance in 14th-century Italian city-states spread risk and helped make what became today's global trading network viable. The loss of a vessel was no longer an existential disaster for its owner. Provided insurers understood and priced the risks correctly, the market could grow in lockstep with mercantile activity. Data resilience has a similar dynamic, although it is rarely described in those terms. Modern IT would be impossible without it, and moving off-premises amounts to sharing some of that risk with an outside provider. Risk can evolve rapidly. The Lloyd's of London insurance market prospered around the turn of the 20th century amid two technology booms: steam-powered merchant shipping and the cable and wireless networks that supplied the information needed to coordinate global trade. Then war came. Insurers responded by separating war risk into its own category, with government support helping to keep ordinary marine insurance affordable while covering higher-risk voyages separately. It may take the combination of geopolitical instability and intense competition for storage to make enterprises perform similar calculations about their own precious cargoes of information. When AWS can lose an entire region and critical national infrastructure can fail without an adequate backup, those risks have plainly not been priced in properly. There is no global market where AWS, NATS, or your organization can buy the data equivalent of war-risk insurance, although imagining what one might look like suggests some intriguing possibilities. Until something like that happens, we'll have to play by the old rules. Work through what happens if your primary data store disappears. Match backup provision to the actual risks, and if you cannot afford to protect all the data your organization needs to survive, determine how to survive with less. Backups, like insurance, are all too easy to let slide. Then a drone sinks your ship. You can't say you weren't warned. (R)
External Content
Source RSS or Atom Feed
Feed Location http://www.theregister.co.uk/headlines.atom
Feed Title www.theregister.com - Articles
Feed Link https://www.theregister.com/
Reply 0 comments