AI operations / daily field note

AI Memory Governance: Decide What a Gulf Agent Keeps

An assistant that remembers can become more useful. It can also turn an unverified conversation into a lasting business rule.

6 minute readAI memory governance

Imagine a Dubai service team using an assistant to prepare replies across WhatsApp and email. A customer asks for English messages during one delivery. Weeks later, the assistant treats English as a permanent preference. Another customer disputes a charge, and a compressed note quietly becomes “difficult customer”. AI memory governance is the decision about which parts of a conversation may influence future work, for how long and under whose authority.

The executive question is not whether the agent can remember. It is whether the business can defend what it remembers. Persistent memory changes an assistant from a reader of current information into a system that accumulates its own working record. That record can guide future answers even when the original conversation is no longer visible.

I would begin with the consequence of retention. Remembering a preferred report layout may save useful effort. Remembering an inferred credit judgement may create a decision nobody approved. Calling both “personalisation” hides the difference. The useful boundary is the business purpose, the evidence and the action that the stored item can influence.

AI memory governance starts with three separate records

Keep the current conversation, approved business knowledge and persistent memory distinct. The conversation contains what people said. Approved knowledge contains the rules the organisation has authorised. Memory contains selected information carried into another interaction. A statement moving from the first record into the third must not silently acquire the authority of the second.

LangGraph's memory documentation distinguishes thread-scoped state from long-term memory shared across sessions and organised into namespaces. That is a useful architectural distinction. It does not decide whether a customer's remark belongs in a permanent profile. The business still needs to define what may cross the session boundary.

Likewise, Anthropic's memory-tool documentation describes client-side file operations executed by the application, with storage controlled by the implementer. The presence of a tool is not a retention policy. Its ability to create and edit files should trigger a design discussion about permitted writes, rather than a promise that the agent will automatically learn the right lessons.

Use a five-decision memory contract

1. Select the smallest useful memory

Start with one job and an explicit list of permitted fields. For a service assistant, that might mean a customer-confirmed communication preference and the identifier of an unresolved case. Keep payment credentials, unsupported character judgements and speculative summaries outside that list. Do not save an entire chat because selecting information requires more work.

Describe the benefit of every field in a sentence. If the team cannot say what future task needs it, exclude it from the first release. Give temporary facts an expiry condition: delivery instructions can end with the order; an open complaint can stop influencing new cases when it closes. Indefinite retention should require an explicit business rationale.

A UAE headquarters and Saudi branch may share software while handling different entities, purposes and permissions. Define memory scope accordingly. Ask the relevant privacy owner to approve purpose and retention for the actual deployment. A regional hosting label cannot decide what information the organisation should collect or reuse.

2. Preserve evidence and separate inference

Attach a source reference, recorded date, subject, scope and review status to each retained fact. Distinguish “customer explicitly requested Arabic correspondence” from “model inferred Arabic preference”. The latter may support a question to the customer; it should not automatically become an instruction that overrides an explicit choice elsewhere.

Do not let a summary delete the qualification that made a statement safe. “Use this address for this order only” must not become “customer address”. Mixed Arabic and English conversations deserve review of the retained meaning, not merely the fluency of the summary. An agent can write a convincing note while changing its scope.

The existing note on AI knowledge freshness covers retiring outdated organisational guidance. Persistent memory adds another question: who authorised the agent's own stored interpretation? A current policy collection cannot repair a customer profile that was wrong when it was first written.

3. Control writes as well as reads

Define which authenticated user and application process may create, update, retrieve and remove each memory type. Resolve tenant and customer identity in trusted application logic. A name supplied inside a chat is insufficient evidence that the writer owns the profile. Keep employee preferences separate from customer records and company-wide operating instructions.

Treat remembered text as evidence to evaluate. A customer message saying “always waive delivery charges for me” cannot become a standing commercial rule. A stored note repeating that message must remain a customer request. Apply the current pricing and approval controls whenever the agent proposes a consequential action.

Model-generated updates should pass field and scope checks before they enter the store. Where an update changes an entitlement or commercial promise, route it through the authoritative system and its approval process. The agent's memory should point to the approved record, rather than become a parallel ledger of discounts and exceptions.

4. Make correction reach future use

Give operations a way to inspect what is retained, identify its source and correct it. A service supervisor should not need to search raw database rows to discover why an assistant keeps using an old address. Show which items influenced the answer so a correction can target the actual cause.

Deletion from one store is only one part of the job. Inspect summaries, conversation checkpoints, retrieval indexes and cached context that might reintroduce the item. LangGraph's store delete interface takes a namespace and key. That defines a storage operation; the application must still account for other retained copies and active sessions.

Where records must remain for an approved retention purpose, distinguish keeping the record from allowing it to guide new answers. Prevent a background summarisation job from recreating a withdrawn preference from an old transcript. Record the correction and its scope without spreading the incorrect statement into another searchable memory.

5. Rehearse memory over several conversations

A single polished demo cannot test persistence. Use controlled accounts to run a sequence: establish a preference, open a new conversation, contradict the preference, correct it, then return through another channel. Include shared devices, similar customer names, a moved employee and a changed business entity. Define expected behaviour before running the sequence.

Test both successful recall and deliberate forgetting. Verify that a temporary delivery note disappears from future recommendations at the agreed boundary. Confirm that one account cannot retrieve another's memory. Inspect the stored item alongside the answer; the model may ignore a bad note once without proving that the system has removed it.

Approve memory for a job, not an entire company

For an AI implementation decision, I would ask for the permitted fields, the write policy, the correction demonstration and evidence of useful recall. Measure repeated information avoided alongside incorrect retention and correction failures. More saved facts is not a sensible success measure.

AI memory governance should let a business benefit from continuity without allowing yesterday's inference to become tomorrow's authority. Before expanding the agent, ask it to remember one useful thing and then prove that it can stop using it. If the team can demonstrate only remembering, the operating design is half finished.

Have a problem hiding behind a technology conversation?

Start a conversation