Downtime in IT is more frequent: either the problems evolve faster than the solutions or companies invest in a poor approach to resilience.
DFrom a pragmatic point of view, architectural complexity is not getting any lighter: companies tend to add additional monitoring tools rather than ensuring they have unified visibility and context covering all areas. Previously, observability was not a major area of investment; ITOps and engineering devoted the majority of their budgets to services clouddata center modernization and disaster recovery. Now, this trend has been reversed and often even takes precedence over spending on the infrastructure itself.
A sign that approaches are evolving, automation is the second priority for technology executives. By combining observability and automation, organizations can increase the resilience of their operations and reduce the amount of manual effort.
AI tools transform downtime
With its ability to automate complex tasks, predict outages, and dramatically reduce response times, businesses are turning to AI to find a solution to outages and service degradation. This technology is emerging as a powerful ally for digital resilience, with spending averaging up to $24.5 million per year on AI tools designed to prevent and manage service interruptions.
According to a recent study, AI is adopted by 44% of security, IT operations and engineering professionals. It allows them to tackle two major problems: the slowness of investigations, which prolongs the impact on the activity, and the inability to locate the true source of an interruption, some interruptions being too brief to justify an investigation. It also supports teams lacking visibility into external dependencies or overwhelmed with alerts, allowing them to only have to focus on the most serious incidents.
On a more granular level, generative AI assistants help human investigators quickly correlate data, summarize complex incidents, and suggest a course of action to accelerate analysis. AI agents go even further: they are able to autonomously diagnose problems and apply common corrective actions, such as reverting to an earlier version of code, and will subject any critical interventions to human approval.
The paradox of AI, essential human supervision
While its effectiveness is gaining ground and proven, companies tend to hastily deploy AI systems without the necessary safeguards and oversight, while it is essential to clearly define responsibilities and escalation procedures. Quite simply because this technology is proving to be a double-edged sword: technology leaders admit that AI has already caused some form of downtime in their business. Another worrying fact is that the growing use of unapproved AI tools (“shadow AI”) in their work poses a direct threat to the governancecontrol and integrity of a company’s security operations.
In addition, the risk of seeing generative AI passing into the hands of cybercriminals is high. Many companies have already experienced malicious code injections or data poisoning incidents. This is proof that bad actors are actively probing AI systems for weaknesses, revealing critical flaws in the security of the infrastructure in the process.
AI has become a powerful ally in the fight against downtime. But without close control, it can create new risks. It is important to systematically combine the speed of AI analysis with the judgment and supervision of human experts. Any corrective measures or automated actions must be validated and must go hand in hand with strong governance frameworks to avoid clandestine AI and secure AI-driven workflows. The future of digital resilience lies in this collaboration between humans and agents, where AI serves the expert, not the other way around.