Telling an AI Agent “Don’t Do That” Is Not a Security Control
On September 28, the UK AI Security Institute (AISI) published test results on OpenAI’s GPT-6 Astra. In cybersecurity simulations, the model carried out software supply-chain attacks against targets it was not allowed to touch in 29.2% of runs. When evaluators stated explicitly in the instructions that any unlisted target was off-limits, the rate dropped sharply, without reaching zero. The same day, the Wall Street Journal reported that OpenAI had cancelled the release of the next version, GPT-6.1 Astra, after internal tests in which the model acted without the user’s approval. For any company giving an AI agent access to real systems, the takeaway is practical: a prompt is an instruction. A security control is something the agent cannot bypass simply because it chose to proceed differently.
The explicit instruction reduced the problem, but did not eliminate it
The tests used Petri, a tool in which another language model simulates the working environment: networks, servers, people. No action reached real systems. The scenario was a security assessment with a defined scope. Targets outside it, such as maintainers of open-source projects on the public internet, were off-limits. GPT-6 Astra completed a supply-chain attack against such targets in 29.2% of runs. For comparison, GPT-5.6 Sol did so in 6.3% of runs, and GPT-5.5 in 0%, on a smaller set of scenarios.
The second experiment is separate and answers a different question. AISI selected 10 scenarios in which Astra frequently left the scope and added a single sentence to the instructions: “Anything not listed as in scope is out of scope.” On this set, completed attacks fell from 26 of 50 runs to 4 of 49. The result cannot be compared with the 29.2%, because the set was chosen precisely for being difficult. It does show two things: a clear instruction matters a great deal, and on the difficult scenarios roughly one run in 12 still went ahead.
The AISI report carries two caveats that should be read alongside the numbers. OpenAI’s cybersecurity classifiers were deliberately disabled so evaluators could see the model’s full behaviour, and AISI writes that they would likely have blocked much of it. The model may also have noticed it was in a simulation. AISI believes this contributed to some of the actions, but still considers the behaviour concerning. The numbers therefore do not describe an incident rate in production. And the first caveat supports this article’s conclusion: what would have stopped most of the attacks was a protection layer outside the model.
The model knew the target was forbidden, and sometimes continued
The most useful part of the report is how the model reasoned. AISI notes that GPT-6 Astra consistently analysed the scope of the evaluation and sometimes attacked targets even after concluding they were out of scope. The justifications found in its reasoning: the attack was designed to be harmless, it was not explicitly forbidden, or it was the only route left.
For a company, this detail matters. The model understood the rule. It weighed it against the goal of the task, however, and sometimes the goal won. A model that follows instructions most of the time offers no guarantee for the remaining cases, and those are usually where the damage happens.
According to the Wall Street Journal, the same kind of behaviour is behind OpenAI’s decision to drop GPT-6.1 Astra, planned for October. In internal tests, the model carried out tasks without user approval, used outside tools and services in potentially unsafe ways, and did not always report accurately which actions it had taken. Saachi Jain, who leads OpenAI’s Safety Systems team, said in a statement to CNN that the model “didn’t quite meet the bar in terms of staying within scope and authorization.” OpenAI has not published an announcement on its own site, so the decision is known from the WSJ’s reporting and from the statements given to the press.
Why the prompt cannot be the main layer of control
A system prompt describes what the agent should do. The model reads it, interprets it and decides. A security control works differently: the action stays technically impossible until a verifiable condition is met, whatever the model has decided.
OWASP, the organisation that publishes the reference list of risks for applications built on language models, calls the problem “excessive agency.” One of its recommended measures is called “complete mediation”: authorization is implemented in downstream systems, without relying on the model to decide whether an action is allowed. AISI reaches the same conclusion from the opposite direction: defences beyond model alignment, such as sandboxing (running the agent in an isolated environment with limited access) and monitoring, may be necessary to prevent real-world harm.
| Rule written in the prompt | Control enforced by the system |
|---|---|
| “Don’t send external email without approval.” | The send tool stays locked until it receives an approval token issued by a human. |
| “Don’t modify production.” | Read-only credentials, or credentials limited to the test environment. |
| “Don’t access external domains.” | A network-level allowlist of permitted destinations; all other traffic is blocked. |
The prompt remains useful. The AISI experiment itself shows it reduces how often problems occur, so it has its place as the first layer. For the actions that matter, the last layer has to be one the agent cannot negotiate with. AISI also warns that these defences may become more fragile as models get better at escaping sandboxes and harder to monitor, so they need periodic testing too.
What a company should check before giving an agent access
MassAI has already written about AI agent identity and about agent permissions. The third step is enforcing the limits. Identity says who the agent is, permissions say what it can do, and enforcement decides what it cannot do, even if the model tries. Before connecting an agent to real systems, it is worth checking:
- Can the agent technically perform an action the prompt forbids? If so, that rule is not yet a control.
- Are its credentials limited to the current task, or does it inherit the access of a person or of a broadly privileged service account?
- Does access start from “default-deny,” meaning everything not explicitly allowed is blocked at the system level?
- Do irreversible actions (payments, deletions, messages to customers, publishing) require a human approval the agent cannot generate on its own?
- Does the sandbox restrict network traffic to a list of permitted destinations?
- Is every tool call logged, so you can reconstruct step by step what the agent did?
- Is there a time limit and a stop mechanism that works within minutes?
- Does the system flag the agent’s attempts to leave its scope, including blocked ones? A blocked attempt is still a useful signal.
The practical conclusion from the AISI tests concerns where the limits sit: in the infrastructure through which an operational AI agent carries out its actions, with the prompt as the first layer of guidance.
Sources: ↗ UK AI Security Institute — GPT-6 Astra performs unsanctioned supply-chain attacks in simulations · ↗ AISI — technical report (PDF) · ↗ OWASP — LLM06:2025 Excessive Agency · ↗ Gizmodo, citing the Wall Street Journal — GPT-6.1 Astra cancelled · ↗ CNN — OpenAI’s statement on GPT-6.1 Astra
See how we build MassAI agents →