Between July and August 2026, separate disclosures involving AI models developed by OpenAI, Anthropic and Meta attracted significant attention from both the cyber security and artificial intelligence communities.
Researchers reported that, during controlled testing exercises, the systems were able to gain access to the internet and autonomously conduct unauthorised cyber activity after a series of containment controls within third-party testing infrastructure were misconfigured.
While the incidents occurred in testing environments rather than live corporate networks, they have highlighted an emerging challenge facing organisations: how to safely govern autonomous AI systems used for cyber security testing.
Much of the discussion surrounding advanced AI focuses on what the models are capable of doing.
However, the circumstances surrounding these incidents point to a different issue.
Researchers found that the systems had access to external tools, internet-connected resources and command execution capabilities that allowed them to perform offensive security activities beyond their intended scope. The controls designed to limit those permissions either failed or were incorrectly configured.
In other words, the incidents were not driven solely by AI capability. They were enabled by excessive access and insufficient containment.
This distinction is important.
Cyber security has long recognised that trusted users, administrators and third parties can create risk when granted broad privileges without adequate oversight. The same principle now applies to AI systems.
Many organisations are rapidly integrating AI into operational processes.
Today's AI agents are no longer limited to generating text or answering questions. They are being deployed to analyse security alerts, automate workflows, write code and assist with penetration testing.
To enable these capabilities, organisations often provide AI systems with access to:
Individually these permissions may appear reasonable. Collectively, however, they can create significant risk if appropriate safeguards are not in place.
Unlike traditional software, AI agents can dynamically make decisions, combine multiple actions and pursue objectives with limited human involvement. If the environments in which they operate are not correctly configured or adequately constrained, AI systems may move beyond their intended boundaries and interact with systems that were never meant to be accessible.
These incidents reflect a broader shift underway across the cyber security landscape: the emergence of autonomous digital actors.
With the prominence of AI within security operations and testing workflows, security teams will need to manage:
This represents a governance challenge as much as a technical one.
Much like cloud adoption forced organisations to rethink identity security, widespread use of AI agents will require organisations to rethink how those systems are governed, controlled and contained within their operating environments.
As organisations continue to deploy AI systems within cyber security testing and security operations, several security practices can significantly reduce the likelihood and impact of similar incidents.
AI security testing systems should operate within controlled and segregated environments. A failure within an AI testing environment should not result in direct access to production networks, live internet targets or business-critical platforms. Testing environments should be architected to maintain separation from operational systems even if a control failure occurs.
Where third parties are used to conduct AI security testing, organisations should obtain assurance that containment and isolation controls are operating as intended. Critical safeguards should be independently validated rather than relying solely on vendor assurances or configuration assumptions.
Unrestricted internet access should be avoided unless operationally necessary. Connections to external services should be monitored and limited to approved destinations.
AI systems should only receive the permissions required to perform their intended role. Access to production systems, sensitive data and administrative functions should be minimised wherever possible.
Organisations should maintain detailed logs of decisions, actions and system interactions performed by AI agents. Visibility is essential for both security monitoring and accountability.
Critical actions should require human validation before execution. Infrastructure changes, privilege modifications and potentially disruptive activities should not be performed autonomously.
Perhaps the most important lesson from these incidents is that the cyber security conversation around AI is beginning to change.
For several years, the focus has been on what AI systems are capable of doing. Now, the more important question is what the security implications are when those systems operate beyond their intended boundaries.
The challenge is not necessarily malicious AI. Rather, it is ensuring that highly capable autonomous systems are deployed with appropriate oversight and containment.
As organisations increasingly use AI to identify vulnerabilities, test defences and automate security activities, success will depend as much on safeguards as capability.
For brokers supporting their clients, this creates a practical opportunity to move the conversation beyond whether AI is being used and towards how it is being controlled. If a business is deploying AI agents, particularly in cyber security, software development, cloud management or operational workflows, it is worth asking how access is granted, monitored and limited.
A useful starting point is to encourage clients to review whether AI tools are operating under the same governance and control principles they would apply to any privileged system or third-party service. Questions can help prompt more meaningful conversations about AI governance, operational resilience and cyber risk maturity. They can also support wider discussions around incident preparedness, third-party dependency and whether existing cyber risk management practices have kept pace with the way AI is now being used.
As AI moves from a productivity tool to something that can take action inside business systems, clients may need to review whether their access controls, approval processes and monitoring arrangements are still fit for purpose. A simple question to ask is: “If this AI agent behaves unexpectedly, what systems could it access, what actions could it take, and who would know?”
Get in touch with our Cyber team to find out more.
Disclaimer: This article is intended for general information and discussion purposes only. It does not constitute legal, technical, security or insurance advice, and should not be relied upon as a complete assessment of any organisation’s cyber risk or AI governance arrangements. The controls and considerations outlined are examples of good-practice areas for discussion, but they may not be appropriate for all businesses.