A Dubai group gives an AI assistant access to a shared procurement inbox. The first job sounds harmless: read supplier messages, find the purchase-order reference and prepare a reply. Then somebody proposes letting it update the ERP when the details look complete. AI prompt injection defence becomes a business control at that moment, because a supplier email, attachment or linked page can contain instructions written for the assistant rather than the employee.
A better system prompt is useful. It is not an authority boundary. The model still receives trusted instructions and untrusted business content through the same language channel. The executive question is therefore not, “Can the model spot every malicious sentence?” It is, “What is the worst action this workflow can complete when one sentence gets through?”
I would design from that consequence backwards. Assume some hostile or simply confusing content will survive the filter. Make the system prove its source, its permitted task and the business rule for the proposed action before anything changes.
AI prompt injection defence starts outside the prompt
OWASP’s LLM01 guidance says prompt injection can change model behaviour, and that RAG or fine-tuning do not fully remove the vulnerability. It also says foolproof prevention is unclear. That is the right planning assumption: detection reduces likelihood; architecture limits impact.
Draw the complete path. An external email enters a mailbox. A parser extracts its body and attachments. The model sees that material beside operating instructions. It proposes a supplier match, reads an ERP record and may call a tool. Each transition is a trust boundary. Labelling the email “context” does not stop the model treating a sentence inside it as a command.
Separate three roles. The reader may inspect untrusted content. The planner may propose a structured task. The actor may use only a narrow, validated capability. They do not need the same data or permissions. If the reader can both absorb an arbitrary attachment and initiate a payment change, the system has collapsed two very different risks into one convenient demo.
Use five gates before an agent can act
1. Source. Record where every piece of context came from: employee request, customer message, supplier document, retrieved policy or tool response. Preserve the original. External content remains untrusted even when it arrives through an approved mailbox or knowledge base.
2. Task. Translate the user’s request into a small allowed job. “Process the inbox” is not a job. “Extract the PO number and draft a reply for review” is. Bind the task to the requesting identity, current record and expiry time so an old instruction cannot silently become standing authority.
3. Capability. Give the agent the smallest tool and data scope needed for that job. A draft-response tool should not send. A supplier lookup should not edit bank details. A CRM search should not expose every region when the user is responsible for one market.
4. Decision. Validate the proposed action with deterministic rules outside the model. Check identifiers, allowed recipients, record state, value thresholds and separation-of-duties requirements. Ask for human approval where the consequence is material, unusual or irreversible. Show the approver the source and exact action, not a confident summary of them.
5. Evidence. Store the request, source references, policy result, tool parameters, approver and outcome. Redact sensitive content where appropriate, but keep enough evidence to reconstruct the decision. Monitoring a model response without the resulting business action leaves half the incident invisible.
The NIST NCCoE concept paper on software and AI agent identity frames the unresolved questions clearly: agent identity, least privilege, authority for a specific action, delegation, human authorisation and auditable intent. It is a concept paper, not a finished control standard. Still, those questions make a useful design review.
Test the workflow, not a list of jailbreak phrases
A red-team prompt copied into a chat box proves very little about an operational agent. Build tests from the content the workflow actually reads: an email footer, PDF, spreadsheet cell, product description, CRM note, website and tool output. Include hidden markup, quoted threads, multilingual content and instructions split across several files.
For every test, record four outcomes. Did the system recognise the untrusted instruction? Did the model’s plan drift from the authorised task? Did a policy stop the proposed tool call? Could any data or external state change before the stop? The last two matter most.
Microsoft’s guidance on indirect prompt injection recommends defence in depth, combining probabilistic and deterministic mitigations, isolating untrusted content, monitoring plan drift and using human verification for risky actions. No single prompt shield should carry the whole control story.
Run the workflow first with read-only access. Then allow proposals without execution. Add one reversible write only after the team can see rejected attempts, false positives and the evidence an operator receives. This staged permission path is more informative than a broad launch followed by a policy document.
The adjacent note on agentic AI permission budgets explains how to price authority by consequence. The guide to AI incident response covers the stop rule once behaviour becomes unsafe. Together they turn a security concern into an operating design.
Make the stop operational
Define who can remove tool access, revoke the agent identity, quarantine a source and pause a workflow. Test those controls during normal operations. A kill switch hidden behind the same compromised assistant is decoration.
Watch for new tools, broader scopes, unexpected destinations, repeated policy refusals and actions that no longer match the original request. Reassess when a workflow starts reading a new content type or writing to a new system. The risk changed even if the model did not.
This is part of practical AI strategy and implementation in Dubai: decide where the model may interpret, where code must enforce, and where a person remains accountable.
AI prompt injection defence is credible when one poisoned document can waste a model call or trigger a review, but cannot borrow the organisation’s authority. Do not ask the demo to promise perfect obedience. Ask it to show the exact gate that blocks a plausible wrong action. If that gate does not exist outside the prompt, the agent is not ready to act.