After incidents involving AI agents, are companies facing a new kind of cybersecurity threat?


In recent weeks, four separate disclosures involving OpenAI, Anthropic, Meta and the UK’s AI Security Institute (AISI) have raised fresh questions about how AI agents are tested before deployment.The AISI, a research organisation within the Department for Science, Innovation and Technology that evaluates the safety and capabilities of frontier AI models, disclosed on Tuesday (August 4) that AI agents powered by Anthropic’s experimental Mythos 5 and OpenAI’s flagship GPT-5.6-Sol had engaged in unauthorised actions during cybersecurity evaluations designed to assess their capabilities.Responding to the findings in a statement to The Indian Express, Britain’s AI Minister Kanishka Narayan said the actions had been detected during “routine cybersecurity testing” and that “AISI caught it and stopped it quickly”. He described identifying “new behaviour like this” and sharing the findings as “exactly what AISI was set up to do”. Cybersecurity breaches at OpenAI, Anthropic and Meta.The incidents raise a question: Do these indicate a new class of cybersecurity risk, or are they primarily revealing the limitations of how increasingly autonomous AI systems pursue goals during testing? What are agents and why do they need evaluations? Unlike chatbots, or by extension Large Language Models (LLMs), which provide answers in response to a prompt or question, AI agents enjoy greater autonomy and are designed to pursue goals. Beyond text generation, agents may engage in complex tasks such as reading and sorting email or analysing financial data. These tasks require them to make decisions, choose their own sequence of actions, and interact with external systems. This autonomy makes their behaviour harder to predict, making evaluations that simulate real-world scenarios increasingly important. Such evaluations provide an opportunity for decision-makers to anticipate unexpected behaviour and course-correct before deployment. How might AI agents pose a risk? Unlike traditional software, AI agents are increasingly given the authority to act on a user’s behalf, whether by accessing email, browsing the web, writing code or interacting with other software. That means errors, unexpected behaviour or manipulation can have real-world consequences rather than remaining confined to a conversation.Story continues below this ad Researchers have broadly identified four stages at which these risks arise. A 2025 paper titled AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways groups them into four categories spanning the information an agent receives, the way it reasons, the actions it takes using external tools, and its interactions with websites, software services and other AI agents. At the input stage, attackers may use prompt injections—hidden instructions embedded in web pages, documents or other content an AI agent processes—to manipulate what the agent sees or does. During reasoning, flaws in planning or decision-making may cause an agent to pursue unintended objectives. When using external tools, excessive permissions or compromised software can lead to unintended actions, such as sending emails or modifying code.Story continues below this ad Agents that interact with websites, software services or other AI agents introduce risks that can spread across connected systems rather than remaining confined to a single application. Do these incidents indicate a cybersecurity risk? Until recently, cybersecurity largely meant defending computer systems against human adversaries. Think cybercriminals stealing bank details, ransomware gangs locking hospital networks, or state-backed hackers targeting government systems. The threat actor was a human with intent and control, and AI was a tool they might use. AI agents complicate that picture. Unlike chatbots that simply respond to prompts, agents are designed to pursue goals independently, using browsers, coding tools and other software to complete tasks. Recent evaluation incidents show
Discussion (0)