The recent cyber security breaches involving Hugging Face, Anthropic and Meta have reignited concerns about the risks posed by increasingly autonomous AI systems. What began as isolated reports of AI models acting beyond expectations quickly became a broader industry conversation about governance, accountability and control.
OpenAI revealed that one of its systems compromised external infrastructure during evaluation, including Hugging Face. Anthropic disclosed that some of its models interacted with real-world systems during testing where this should not have happened. Meta confirmed that one of its most capable agentic coding models exploited a vulnerability in a third-party service after inadvertently gaining internet access during an evaluation.
Perhaps most concerning was a report from the UK’s AI Security Institute, which documented 19 instances of AI agents taking autonomous, unsanctioned actions against real people and organisations, including a case where a model created fake online identities to socially engineer an open-source maintainer into approving malicious code.
For security leaders, these incidents matter because they show AI systems are increasingly capable of acting like attackers, even when that was never the intended outcome.
Access, not inventory, is the real risk
After every incident, many organisations rush to count how many AI agents they are running or where they are, but this is the smaller part of the problem.
The reality is that no inventory will ever be complete. Agents can spin up, duplicate themselves and disappear within minutes. By the time you’ve catalogued one batch, another hundred may already exist. As adoption accelerates across the APAC region, a complete inventory isn’t just difficult; it’s unrealistic.
Delinea’s latest Identity Security Report shows that nearly 90% of Australian organisations report at least one identity visibility gap, and more than half say those gaps are most likely to persist in AI-related environments. Since organisations will never have full visibility into every identity, privilege and access path an agent touches, the only workable strategy is to limit the damage when one inevitably reaches something it shouldn’t.
This is where Anthropic’s disclosure is instructive. The model wasn’t acting maliciously in a human sense; it was pursuing an objective using the permissions and connectivity it had been given. This means we need to treat AI agents as a new class of privileged identity, governed and enforced continuously at the point of action rather than reviewed after the fact, and the exact agent count stops mattering.
The accountability gap
When a human breaches a system without authorisation, accountability is relatively straightforward — legal and regulatory processes exist to establish who is responsible. When an AI system performs the same action, responsibility is far murkier. Yet ambiguity cannot become an excuse.
Our research found 80% of Australian organisations can’t consistently explain why a non-human identity performed a privileged action. If you can’t explain why an action happened, what privileges were used, or whether it stayed within acceptable limits, you have no way to manage the risk, let alone answer to a regulator or customer afterwards.
Compounding this, around half of regional organisations report having no viable alternative to standing privileged access for non-human identities, meaning access to sensitive systems often remains open long after the task that required it is finished. Guardrails inside the model are no substitute for this kind of control.
Recent testing showed advanced agents attempting deception and social engineering despite built-in safeguards, proof that oversight has to sit outside the model, observing behaviour and enforcing policy independent of what the AI itself decides is appropriate.
As AI systems take on more autonomous roles, Australian and APAC organisations will increasingly be expected to prove, not just claim, that appropriate controls were in place, which means knowing what an agent did, what access it used, and when someone intervened.
That requires traceability built into the infrastructure, not policy statements sitting in a drawer. The recent incidents involving Hugging Face, Anthropic and Meta all point to the same conclusion: AI agents will keep acting in unexpected ways. The organisations that fare best won’t be the ones with the most complete agent inventory. They’ll be the ones that never gave an agent more access than the task required in the first place.

Techedge AI is a niche publication dedicated to keeping its audience at the forefront of the rapidly evolving AI technology landscape. With a sharp focus on emerging trends, groundbreaking innovations, and expert insights, we cover everything from C-suite interviews and industry news to in-depth articles, podcasts, press releases, and guest posts. Join us as we explore the AI technologies shaping tomorrow’s world.









