Key takeaways
- Prompt injection attacks: Cyber attackers use prompt injection to exploit a structural blindspot where AI frameworks treat data and instructions identically, preventing models from distinguishing between benign user inputs and hostile, hidden commands.
- The ‘Lethal Trifecta’: Vulnerabilities are multiplied by the intersection of three elements: untrusted input, the sensitivity of data (such as in bookkeeping or CRM systems), and the agent's ability to act externally from the enterprise environment.
- ‘Zero-click’ threats: The severity of this risk is highlighted by ‘zero-click’ attacks, such as the ‘Echoleak’ incident, where no human involvement was necessary for the AI to bypass security protocols and exfiltrate confidential data from platforms like SharePoint and Teams.
As more companies increasingly give AI tools access to emails and internal systems, they unwittingly introduce a new cyber vulnerability.
Criminals can place hidden instructions within the data that AI interacts with; these are ‘prompt injections’. Simply put, prompt injection attacks enable criminals to co-opt and manipulate large language models or agents into executing the attacker's instructions over the pre-arranged tasks set by the enterprise system. Those AI frameworks are then vulnerable to leaking data, bypassing system cybersecurity protocols and defences, or to being the conduit for some other unauthorised activity.
No Human accessory needed
In 2024, a SlackAI user accessed information via a public Slack channel that included hidden prompts to exfiltrate confidential company information. On clicking into the channel, the AI assistant retrieved and encoded stolen data, including it alongside the publicly available information the user had requested.
The stakes were decidedly higher when a Microsoft 365 Copilot account was co-opted by a prompt hidden in a phishing email. Known as ‘scope violation’, the hidden instructions led Copilot to ignore its own permissions and resrictions. This left confidential data across Sharepoint, OneNote and Teams open to theft.
Nicknamed ‘Echoleak’, this was the first widely disclosed ‘zero-click’ prompt injection attack: no human involvement was necessary to deliver its malign payload into the enterprise system.
The appeal to criminals
The threat is serious enough for the UK’s National Cyber Security Agency and international allies to pronounce prompt injection as “the most persistent and difficult-to-fix threat” facing these agentic AI systems.
Prompt injections are increasingly appealing to cyber criminals because of their economic viability. Attackers only need one prompt to work. They have a massive attack surface area and a multitude of entry options do be able to do this. For example, prompt injections can be introduced via other agents’ instructions or actions, external databases, webpages and spoof websites, unsecured documents, or phishing emails, as in the case of Echoleak.
The threat is magnified through the autonomy of agents and the potential enterprise system access points that they have. Recent research commissioned by IBM found AI-led breaches in particular surged 56% globally over the last year, costing an average of $6m per breach. It was the second most costly form of cyber breach, just under model inversion attacks, which reverse engineer an AI or machine learning model’s outputs to access sensitive training data.
According to the research, more than a quarter of organisations globally and 22% in the UK, reported AI-generated attacks.
It’s all just data to AI
Prompt injections exploit a particular structural blindspot of AI frameworks: large language models treat data and instructions – both written in natural language text – as one and the same. They cannot discern between either form, let alone between benign and hostile code.
“That's something that even the best models today, Fable and OpenAI GPT 5.6, and the other ones, are facing,” says Ken Bastiaensen, co-founder and CTO of AI-driven accounting systems firm Ravical.
“In technical reports, prompt injection risk is always one of the elements,” he adds “They say they’re getting better [at distinguishing between data inputs] with 90% or 95% accuracy or whatever. But it doesn't matter, because none can claim to exclude errors altogether. They're improving, but that risk is still there.”
Behold the ‘lethal trifecta’
At the heart of that vulnerability is what Ken Bastiaensen calls “the lethal trifecta”: three elements that include:
- untrusted input, where large language models cannot differentiate between human or agentic instruction;
- sensitive data in bookkeeping systems, reporting systems, document management systems or CRM systems, all of which give the agent context and power to be of use; and
- the agent’s ability to act externally of the enterprise.
These, in turn, multiply ways in which a firm's systems and data can be breached.
Bastiaensen advocates a “holistic overview approach to defence” against prompt injection attacks, one that incorporates the traditional, ‘first principles’ alongside specific measures for monitoring each agent and any associated “untrusted input.”
“You have to follow that agent to track its capabilities, inputs, and scope to act. It’s all of these things, there is no “one or the other”. You have to see for each agent what its capabilities are to be able to lock it down.”
The National Cyber Security Centre (NCSC) makes these recommendations to create crucial data boundaries for AI models:
Apply least privilege: give agents the minimum access needed, for the shortest time required.
Limit scope: restrict agent access, actions and when it can take them.
No long-term credentials: use temporary credentials where possible; revoke elevated access when tasks are complete.
Secure defaults: design applications with safe configurations, secure protocols and appropriate validation built in.
Understand dependencies: manage supply-chain risk for third-party components, models, tools and integrations.
Monitor behaviour: scan for unusual or unexpected activity across tools, workflows and connected systems.
‘Red teaming’ deployment: how could the system be misused, manipulated or caused to behave unexpectedly?
Incident planning: response plans to cover agentic AI failures, misuse and loss of control.
Cyber security awareness
Each year ICAEW marks global Cyber Security Awareness month with a series of resources and a podcast addressing the latest issues and how to protect your business.