Every Friday: An AI topic. A clear judgment. No bullshit.
What happened
Palo Alto Networks Unit 42 drilled into the wound twice this week — and did so with surgical precision. On Tuesday, researchers published a study that shows how LLM guardrails can be systematically defeated using a genetic algorithm. The principle is simple and precisely why it is worrying: an automated fuzzer iteratively mutates prompts until the model shows the desired forbidden behavior. Not a single LLM tested was immune. Evasion rates varied between low single digits and high percentages depending on the model and keyword — but the core message remains: It’s a question of effort, not impossibility.
The second study, published today, looks at AI in the offensive — specifically, how attackers are already embedding AI in malware. Unit 42 distinguishes between three categories: AI as the author of malicious code, AI as an autonomous command-and-control decision maker, and local AI agents in malware – the latter still largely theoretical. The specific examples are revealing: An infostealer who invokes GPT-3.5 without gaining anything tactically – Unit 42 aptly calls this “AI Theater”. On the other hand, a Golang dropper that uses LLM to check whether the target system is valuable enough for an infection – that’s “AI-gated execution”, and that’s actually clever.
The macroeconomic context provided by IBM AI is not just a tool for developers – it has also long been a target for attacks.
To summarize this week in one sentence: The foundation on which many AI security promises were built is cracked — and the attackers have drills.
What the world says about it
In the DACH region, the topic is treated with the usual mix of technical sobriety and bureaucratic caution. The BSI has already classified prompt injection as a priority risk in its AI security guidelines – OWASP LLM01:2025 is no stranger there. In its most recent assessment of the situation, the NCSC-CH pointed out the increasing professionalization of attacker groups, without, however, explicitly naming AI as an accelerator. That will have to change. Austria’s CERT.at has so far been remarkably silent on the topic of AI-specific attack vectors.
In the professional public – security researchers, CISOs, red team service providers – the reaction to the Unit 42 studies is not with panic, but with a collective “We told you so.” The argument that LLMs are not safety limits but probability machines has been a consensus among experts for months. What is new is the empirical support through automated, reproducible fuzzing. This gives the argument the clout it needs in boardrooms and budget discussions.
What is neglected in the media reporting – also in Germany and Switzerland – is the threat posed by RAG systems via indirect prompt injection. Anyone who operates an internal company AI system with external data sources today has created an attack vector that is fundamentally different from classic SQL injections – and for which most security teams do not yet have established defense processes.
What we think about it
Let’s start with what’s not surprising: that LLMs are hackable. All software is hackable. But what’s coming into sharper focus this week is the structural miscalibration in the industry. Companies — in Switzerland as well as in Germany and Austria — deploy LLM-based systems with the implicit assumption that the model itself functions as a last line of defense. It doesn’t. It never worked that way. And Palo Alto Networks now has the data to make that point hard to miss.
The “AI Theater” phenomenon in the malware world is actually almost comical: someone builds an infostealer, integrates GPT-3.5 — probably because it sounds better in the pitch deck — and doesn’t achieve a single tactical advantage. The barrier to entry for malicious code is actually lowered with AI code generation, but that doesn’t mean that every script kiddie will suddenly reach APT level. What is increasing is the number of attempts. What remains the same is the quality of the unimaginative attacks.
The “AI-Gated Execution” category is more serious. A dropper that uses LLM to decide whether a target is worth the effort of infection reduces false positives on the attacker side and increases the time to discovery. This is operationally relevant — not as science fiction, but as today’s code that Unit 42 has analyzed. IT security teams in DACH who believe that their company size protects them from targeted attacks should see this example as a wake-up call.
What does this practically mean? Three things. First: Anyone who runs LLMs in productive systems needs layered controls – the model is a layer, not the perimeter. Second: Least privilege for AI agents and RAG systems is not an option, but a requirement. Third: Adversarial testing – i.e. regular, automated fuzzing of your own AI systems – must be included in the CISO budgets for 2027. Anyone who waits for a vendor to solve the problem is waiting for the apocalypse in installments.
By the way, the equation is not just “AI strengthens attackers.” IBM X-Force maintains, and Unit 42 implicitly confirms: Defenders benefit equally from the same tools. Those who use AI-powered threat hunting, automated anomaly detection and adversarial red teaming are faster than ever before. The difference is not in the availability of the tools – it is in the consistency of the application. And the defenders in DACH still have some catching up to do.
The verdict
| criterion | Evaluation |
|---|---|
| substance | ⭐⭐⭐⭐ |
| DACH relevance | High |
| Time frame | Now |
Conclusion in one sentence: LLMs are not firewalls — treating them as such does not have an AI problem, but an architecture problem.
The AI verdict appears every Friday. Subscribe for free:
👉 aisyndicate.ch/#/portal