Opinion

Two years after Crowdstrike, businesses face even greater risks‍

By
By
Kashif Nazir

When Crowdstrike sent out an update on 19th July 2024, it didn't know that it would become one of the most talked about outages in IT and cybersecurity for some time. Millions of Microsoft systems and servers went down simultaneously across public and private sectors. People couldn’t make purchases, as card machines failed to work, flights experienced major disruption and, critically, some hospitals had to revert to backup methods as systems froze.

The cause of the incident was down to a faulty update that, as IBM describes, “triggered an infinite boot cycle of the operating system, leaving the systems unable to boot correctly”. A small change leading to widespread disruption around the world. But while the impact was vast, the remediation by Crowdstrike was pretty swift. And the company itself has since gone from strength to strength again, illustrating the importance of honest self-evaluation.

In response to the outage, assessing and testing vendors more stringently and implementing strategies like staged rollouts were some of the key lessons businesses took from it. But this overlooked the underlying reasons behind outages. The more dangerous threats are the changes that don’t lead to immediate damage, the configuration shifts that build up within an estate, drift and then eventually clash with other changes to cause failures.

The problem is many companies now oversee vast IT estates comprising thousands of endpoints. This means IT teams can have limited or no visibility into many changes until they lead to a damaging event. And as AI, cloud and automation tools continue to evolve, the risks are becoming more apparent. But why?

The dangers of auto-update

Enterprises are continuing to deploy the latest AI tools and cloud services. IT teams are also automating more internal processes, like having an AI copilot perform certain engineering tasks. The range of devices and integrations in a company’s network mean that thousands of small configuration changes can occur in one day alone. Crucially, as with what happened with Crowdstrike, so much of this change is scheduled and pushed by third-party providers, and these can take place at any time. 

So, to mitigate this invisible change, many organisations will have auto-updates on to ensure all devices are immediately secure with the latest security patches. It’s a sensible approach, up to a point. When it came to Crowdstrike, many systems went down because the majority of subscribers had auto-update enabled, showing the risks of this strategy when the update is faulty. 

As a result, experienced admins who have seen these events before will often delay the latest release until they have run n-1 or n-2 – in simpler terms, using someone else’s estate as the test environment. But this also has its downsides, as this method only manages the update versions admins can see. Even by delaying the latest release, there could still be SaaS integrations that update themselves independently, or internal automation tools that correct one setting but disrupt another in the process. It’s these changes and interactions that many of post-Crowdstrike evaluations didn’t consider. 

Where human guardrails fall short

One solution, then, would be more forensic human oversight, which is always important. But an IT team can establish robust control over anything that has been intentionally implemented and still be vulnerable to cyber incidents, as the drift that leads to an outage is usually a deployment that hasn’t been approved. 

Drift is the gap that exists between a system’s actual state and the last known-good baseline it had across its various settings. Most companies will assess drift through a periodic audit or by having a Change Advisory Board (CAB). CABs sign off proposed changes, but as a human needs to submit a change in advance, they fail to meet the demands of a modern estate where the majority of change is automated and conducted by external parties. 

A board meeting can’t govern changes that are unplannable. But the main point is that this isn’t a problem with ‘change risk’ procedures. It’s about trying to manage change that is invisible. 

The greatest lesson from Crowdstrike

Changes are continuously being made by humans, automation tools and vendors, regardless of whether they have been approved. This is a reality that IT teams understand. But the risks now are arguably even greater than they were two years ago. The adoption of AI has progressed significantly and has meant there is even more automation and even more external providers in IT estates. Consequently, there is even more risk of invisible change. 

Therefore, in these environments, what matters is that IT teams can see this change when it does happen. To do this, they need to integrate a platform that can automatically and continuously monitor all systems against a known-good baseline to see what is changing and to identify where configuration drift is happening. This creates a detailed evidence log of events that means every change is recorded and can then be authorised. 

When so much change is out of the IT team's hands, continuous real-time visibility into change, instead of quarterly audit or CAB meetings, is the only way to build a scalable system of control. This is the greatest lesson from Crowdstrike. There is always an element of risk from updates. The difference between them staying as routine, small changes and turning into global outages rests on how well companies are able to see how they interact with the other software, servers and dependencies in their estate. 

Written by
July 20, 2026
Written by
Kashif Nazir