There have been numerous articles discussing the challenges posed by the Crowdstrike code issue. In this blog, we delve into two main aspects: first, the problem itself, and second, the observations related to the supply chain, where Crowdstrike plays a significant role.
With Microsoft continuing to monopolise the majority of the market for Operating Systems (OS) it would have been hard for many of us not to have been impacted by Windows computer screens turning blue a couple of weeks ago. A great deal of focus in the news was drawn to airports at a complete standstill, the NHS being denied access to patients' medical records and some cash points being unserviceable. However, there were very few industries, if any, where the blue screen wipeout did not impact either directly or indirectly. So, what happened?
Falcon, a Crowdstrike sensor product, is designed to stop breaches and various types of attacks including malware. Crowdstrike has explained that its product delivers security content configuration updates to its sensors in two ways:
Sensor content is shipped with its sensor directly.
Rapid Response Content is designed to respond to the changing threat landscape at operational speed.
On Friday the 19th of July 2024, according to Crowdstrike: “The issue on Friday involved a Rapid Response Content update with an undetected error.” A bug caused the issue. Without delving too deeply into the technical aspects of OS technology (let’s face it; to most of us, even in the cyber profession, it is another language) let us consider what caused the issue.
A Microsoft OS has 2 modes:
Kernal (Ring 0):
Has greater control, with visibility of the entire system memory map.
More privileged providing core functionality of the OS it provides, it talks to hardware and devices and manages memory.
User (Ring 1):
The user only sees the parts of the memory map that the Kernal wants it to see.
User application.
Application code never runs in Kernal mode and Kernal code never runs in User mode. So, when the application code crashes, the application crashes, but when Kernal mode crashes, the system crashes. When the Kernal detects a critical failure in the code, it blue screens the OS and a reboot is always required as a minimum response. This occurs for all types of OS; the only difference being the colour of the screen (pink for Mac, black for Linux and Blue for Microsoft).
Falcon is more ‘anti-malware’ than ‘anti-virus’ for the server. It analyses a wide range of application behaviours rather than file definitions to proactively detect new attacks so the code has to be in the Kernal to perform. Because of risks to the OS when running code in Kernal mode, Microsoft introduced ‘Windows Hardware Quality Labs (WHQL) Certification’ to ensure the code has been thoroughly tested by the vendor and passed the Windows hardware lab tests and considered compatible with Windows OS.
The WHQL process provides assurance that code or a driver is trustworthy and Microsoft will issue a certificate that remains valid unless there are any changes. However, as with any other assurance process, certifying new or updated products can’t be instantaneous which in itself can expose a vulnerability such as zero-day attacks that can continue to propagate and spread in the meantime.
Wanting to ensure customers receive the latest protection as soon as new threats emerge, Crowdstrike’s approach for the update that caused the Microsoft blue screen incident was to include Definition Files that are processed by the driver but not actually included with it. It is speculated that the definition files were not merely malware definitions but programmes in their own right for which the driver could execute code in Kernal mode even though the update didn’t go through the WHQL process. Consequently, an unassigned and untrusted code was operating in Kernal mode. But what is the link to supply chain assurance?
It's all very well examining a case study such as this to criticise the faults and draw out the failures if only to identify the lessons, but this is not going to be one of those blogs. In actual fact, from a supply chain perspective, Crowdstrike products appear to have been used regularly by Microsoft with a fairly robust assurance process within the design and development of code. But there are some significant considerations regarding the challenges of assuring the supply chain we can indeed take away.
Any security incident that breaches integrity, availability and/or confidentiality will have damaging impacts to an organisation or business. In this case, prevention of operations rippled far beyond Microsoft to all its global customers across virtually all industries, not to mention the damage to Microsoft’s reputation (and Crowdstrike too). No matter the size of a business or company, with the best security that could possibly be implemented and the most mature assurance processes, there will still be vulnerabilities from third parties.
Securing a supply chain appears the same as managing cyber security risk. This is indeed the approach, but the fundamental difference is being able to identify and maintain visibility of risks beyond a company’s remit. A great deal of risk to supply chains stems from a lack of visibility of third-party decisions and practices and indeed the processes behind them. In this case, it was Crowdstrike’s decision to include Definition Files in their Drivers which effectively ‘sidestepped’ the WHQL assurance process. Providing assurance against such ‘unknown unknowns’ whilst allowing third-party suppliers and vendors to access the information they require to perform their service makes supply chain security a complex challenge where every aspect can often present a degree of risk ……which isn’t discovered until it’s too late, as in this example. A proactive company might analyse multiple scenarios as part of its Risk Management strategy, but there is still the necessity for robust incident management and well-practiced Business Continuity Plans too.
The protection of integrity, availability and confidentiality of products, services and information is essential to a company such as Microsoft and, as we have seen with the blue screen incident, essential to those dependent on their services too. Supply chain security is a multi-faceted problem which requires a coordinated and collaborative approach from all parties. The companies, businesses and organisations that get this right recognise and understand these risks and strive to govern the visibility of the ‘unknown unknowns’ at every tier with analytics and monitoring capabilities.
Nexor provide a supply chain assurance service to our customers, if you are concerned about your supply chain contact us to today to learn more about how we can help you.
References:
1. Forbes: CrowdStrike Reveals New Details About What Caused Windows Outage, Kate O'Flaherty, Senior Contributor, Cybersecurity and privacy journalist, Jul 24, 2024,11:59am EDT, CrowdStrike Reveals New Details About What Caused Windows Outage (forbes.com)
2. YouTube, Dave’s Garage: CrowdStrike IT Outage Explained by a Windows Developer, https://youtu.be/wAzEJxOo1ts