When an LLM’s Guardrails Disrupt Cyber Defense

Systems cannot be secured with black-box LLMs.

The statement is signed by Clément Delangue. It comes after an attack against Hugging Face, of which he is cofounder and CEO.

This attack, driven by an agent-based system, exploited two components of the pipeline processing datasets. On one hand, the script loader. On the other, the configuration files. Taken together, they allowed code to be executed on a worker. Then they achieved node-level access, retrieved credentials, and lateral movement across internal clusters.

Hugging Face said it had no evidence of manipulation of its public models and datasets, nor of its software supply chain. It also stated something else… which resonated much more, and which explains Clément Delangue’s remarks: forensic analysis yielded a “bad surprise.” Regarding the initially used LLM (not specified, but it is understood to be an American model accessible via API), safeguards kicked in and blocked the requests.

Hugging Face says it had to switch to a self-hosted open-weight model. Specifically, GLM 5.2, from the Chinese company Z.ai. The latter did not miss the chance to promise that it would continue to improve its deployment documentation.

Delangue saw in this a further lever to promote open models, which lie at the heart of his company’s activity. In the context we know, added to the rising prominence of Chinese models on the leaderboards, the topic took on a political dimension.

To read in addition:

Who is Moonshot AI, the publisher of Kimi?
Suspension of Claude Mythos: experts call for lifting the restrictions
The first “agentic” ransomware detected in cyberspace
To optimize LLMs, IBM makes them perform divisions
AI: after the hype, the budget

Dawn Liphardt

Dawn Liphardt

I'm Dawn Liphardt, the founder and lead writer of this publication. With a background in philosophy and a deep interest in the social impact of technology, I started this platform to explore how innovation shapes — and sometimes disrupts — the world we live in. My work focuses on critical, human-centered storytelling at the frontier of artificial intelligence and emerging tech.