← Insights

Operations and resilience

The update that grounded flights: change management lessons

15 July 2026 · 8 min read · 3 public sources

A team reviewing performance charts together at a table

On 19 July 2024 a content update from a security vendor took 8.5 million Windows devices out of service. Airlines, hospitals, broadcasters and retailers stopped working within minutes of each other. Delta Air Lines subsequently put its own cost at around $500 million, having cancelled thousands of flights over several days. There was no attacker.

Speed of distribution is itself a risk

Security tooling is deliberately built to update fast, because slow protection is weak protection. The same property means a defective update also arrives fast and everywhere. That trade-off cannot be removed, but it can be shaped: rings of deployment, a canary population that receives changes first, and a defined soak period before general release convert an instant global event into a contained one.

Questions worth asking of every agent you deploy

  • Can we stage this vendor’s updates across rings, or does the product only support all-at-once?
  • Which of our endpoints would we sacrifice as the canary group, and who watches them?
  • If every machine running this agent fails to boot, what is our recovery procedure and how long does it take per device?
  • Which staff can perform that recovery, and does it require physical access or a recovery key held somewhere reachable?
  • What do we tell customers in the first hour, and who is authorised to say it?

Recovery at scale is a people problem

The technical fix in July 2024 was straightforward and the recovery was not, because it had to be applied per machine, often by hand, often requiring a key stored in a system that was itself down. Any continuity plan that assumes remote management will be available has not been tested against the case where the management tooling is the thing that broke. Keep an offline copy of recovery keys and procedures, and rehearse the manual path.

Concentration is a risk register entry

When one product runs on nearly every endpoint in an organisation — or in an entire sector — its failure mode becomes systemic. That does not mean avoiding widely used tools, which are usually widely used for good reasons. It means recording the concentration honestly, knowing which business processes stop if that supplier has a bad day, and deciding in advance which of them need a path that does not depend on it.

The lesson is not about one vendor. It is that the changes most likely to cause a serious outage are the ones considered too routine to review.

Sources and further reading

This article summarizes publicly available research. Source findings retain their original geographic and sector scope.

  1. [01]CrowdStrike outage: we finally know what caused it — and how much it costCNN Business · 2024
  2. [02]Delta grapples with $500M in CrowdStrike outage costsCybersecurity Dive · 2024
  3. [03]What the 2024 CrowdStrike Glitch Can Teach Us About Cyber RiskHarvard Business Review · 2025

Put this thinking to work

Tell us about your estate and we’ll map what to build, secure, or fix first.

Keep reading

More insights