The Knowledge Store Pattern: How to Build AI Agent Memory on AWS, Azure or Google Cloud
Why every serious AI agent needs three kinds of memory, how they fit together, and what to ask your team before you build it on AWS, Azure or Google Cloud.
When an AI assistant gives a wrong or out-of-date answer, the cause is usually what it remembers, not the AI itself. This shows how to give it a memory you can trust, check and afford.
- One store is never enough. Agents ask three different kinds of question, and each needs a differently shaped memory: similarity, relationships and exact records.
- One tier is the truth. The structured records are written first and protected hardest. The vector index and the graph are derived from them and can be rebuilt.
- Memory has to change its mind. Without supersession and retraction, an agent will quote last year’s policy with complete confidence.
- The pattern is cloud-neutral. AWS, Azure and Google Cloud all host it well. The choice turns on where your data, your team and your residency obligations already sit.
Bottom line. Three kinds of memory, one source of truth, and a store that can correct itself, on whichever cloud you already use.
Below: the full article, for technical teams ↓
Watch the episodeAdopt AI Wisely In Your OrganizationThe Software Lens on YouTube →Why agent memory is a leadership problem
When an AI agent gets something wrong in production, the model is rarely the culprit. The cause is usually what the agent was given to read. Three failures come up again and again in the systems we review.
Stale answers. The agent finds a policy that was replaced six months ago and quotes it, because nothing in its memory says the policy was replaced.
No audit trail. A customer, an auditor or a regulator asks why the agent said what it said, and nobody can show what it read before it answered.
Cost that grows faster than value. Everything is embedded into one vector database, the index keeps growing, and the bill grows with it while answer quality stays flat.
All three are design problems in AI agent memory, the part of the system we call the Knowledge Store. Most first-generation designs are retrieval-augmented generation (RAG) over a single vector database: a sensible place to start and a fragile place to stay. The problems are far cheaper to design out at the start than to retrofit after the first incident, which is why this belongs on the architecture agenda and not only on the backlog.
Three questions, three tiers
The instinct, faced with "organisational memory", is to pick one technology and make everything fit. We will embed everything. Relationships are what matter, so use a graph. Just put it in proper tables. Each single choice fails, because agents ask three different kinds of question, and no one storage shape answers all three well.
Figure 1. Three questions an agent asks, and the tier built to answer each
“Have we seen something like this before?”
Finds entries that mean the same thing, even when the words differ.
“What is our current position, and who signed it off?”
Follows the links: what replaced what, who approved it, which team it applies to.
“What exactly does the record say, which version, and on what evidence?”
Every entry with its version, status, owner, approver and evidence. Changes add a new version; nothing is overwritten.
Every write lands in the structured tier first. The vector and graph tiers are built from it, so either can be rebuilt if it is lost.
A vector database is excellent at "something like this", but it cannot tell you which of twenty similar entries is the current one. A graph database is excellent at "what replaced what", but it needs a starting point and is a poor place to keep the full text. Tables hold the exact record but answer neither question at scale. Composed, the three tiers cover each other’s blind spots.
One source of truth, two that can be rebuilt
The single most important decision in the pattern is that only one tier is authoritative. The structured tier is the system of record. The other two are views of it, shaped for a particular kind of question.
That one decision simplifies a great deal for whoever owns the platform. Backup and disaster recovery anchor on one tier, which gets the tightest recovery point objective; the vector index and the graph can be regenerated from it. Audit has one place to look. And there is a clear answer when two tiers disagree: the record wins, and the derived tier is repaired.
It also fixes the order of every write. A new entry is committed to the structured tier first, then embedded into the vector tier, then linked into the graph, then announced to anyone subscribed. If the process stops halfway, the record is still there and the rest is reconciled from it. Reverse the order and the index can end up pointing at entries that were never committed, which no reconciliation can repair. Write order here is not a style choice; it is what keeps the store correct.
How an agent reads: wide, then narrow, then exact
A good read uses all three tiers in turn, each doing the step it is shaped for.
Figure 2. One question, three steps, one answer with its evidence
- 1Find similarVector tier: entries that look related to the question~20 candidates
- 2Keep what is current and permittedGraph tier: drop anything superseded, retracted or outside this agent’s scope~5 remain
- 3Fetch the recordStructured tier: the exact, versioned entriesexact
Counts are illustrative. With sensible indexing the three steps together typically take a few hundred milliseconds.
The three steps cannot be collapsed into one. A similarity search on its own returns entries that look relevant whether or not they are still valid, still in force or allowed for this caller. A graph walk on its own has no good place to start. Used in sequence, each tier covers the step the others cannot.
There is a second benefit that matters as much to a CIO as the answer itself. Because each step is a separate, identifiable query, every read can be logged. When someone later asks why the agent said what it said, you can show exactly what it saw at the moment it decided. That is the audit trail most first-generation agent deployments are missing.
How the store changes its mind
Organisations learn, and what was right last quarter is sometimes wrong today. A memory that cannot change its mind will keep repeating its oldest mistakes. So every entry in the store has a status, and the status moves through a short, controlled lifecycle.
Figure 3. Nothing is deleted; entries change status
Default reads see only validated, current entries. History stays available to anyone investigating a past decision.
The difference between the two end states is worth insisting on. A superseded entry was correct when it was written and has simply been overtaken; it is history, and it may be useful to look back at. A retracted entry was wrong in its foundations, because the evidence did not support it or a compliance point was missed. It is a mistake that must not spread.
Retraction therefore has one more step, and it is the one most often skipped. Other entries may have been derived from the one now retracted. The graph makes it cheap to find them: walk outward along the "derived from" links and queue each descendant for review. Whether that review actually happens is the difference between a memory and a filing cabinet.
Choosing the cloud: AWS, Azure or Google Cloud
Nothing in the pattern is specific to one provider. The three tiers are a design, not a product, and each of the major clouds can host them with managed services.
Figure 4. The same three tiers, mapped to each cloud
AWS
- Structured
- Amazon DynamoDB, or Aurora / RDS for PostgreSQL
- Vector
- Amazon S3 Vectors, or OpenSearch Serverless
- Graph
- Amazon Neptune; Neptune Analytics for Bedrock GraphRAG
- Region in Ireland
- Yes, Europe (Ireland)
Microsoft Azure
- Structured
- Cosmos DB for NoSQL, or Azure Database for PostgreSQL
- Vector
- Azure AI Search, or vector search in Cosmos DB or PostgreSQL
- Graph
- Cosmos DB for Apache Gremlin
- Region in Ireland
- Yes, North Europe
Google Cloud
- Structured
- Spanner, AlloyDB or Firestore
- Vector
- AlloyDB AI (pgvector), Spanner vector search, or Vertex AI Vector Search
- Graph
- Spanner Graph
- Region in Ireland
- No; nearest is London
Each column is a complete Knowledge Store. Pick the one that matches where your data, models and people already are.
On AWS, two recent services change the economics. Amazon S3 Vectors, generally available since December 2025, keeps vectors in S3 itself. AWS puts the saving at up to 90% compared with specialised vector databases, with sub-second queries (around 100 milliseconds when a store is queried often), and it plugs straight into Bedrock Knowledge Bases. It is a strong default for a large store where cost matters more than the last few milliseconds. OpenSearch Serverless remains the better fit when you also need keyword and hybrid search at high query rates. On the graph side, Bedrock Knowledge Bases can now build and query a graph for you on Neptune Analytics, a feature AWS calls GraphRAG, generally available since March 2025. It is useful when the relationships can be extracted from documents rather than written by your own process.
This is not only theory for us. Nexcubator, Aitina Tech’s own platform and the place our research is put to work, uses the AWS-native form of this pattern in one of its components: DynamoDB for the records, OpenSearch Serverless for similarity and Neptune for relationships.
On Azure, Cosmos DB does a lot of the work: its NoSQL API holds the records and supports vector search, and its Gremlin API holds the graph, with Azure AI Search available when search quality is the priority. On Google Cloud, Spanner can hold all three tiers in a single service, with tables, vector search and Spanner Graph side by side. Fewer services means less to operate, which is a real advantage for a small platform team.
Which cloud to choose is rarely about the pattern. It depends on where your data and your team already are, what your models run on, and where the data has to live. For Irish organisations that last point is concrete: AWS and Microsoft Azure both run regions in Ireland, S3 Vectors is available in the Ireland region, and Google Cloud has no cloud region in Ireland, its nearest being London. Making that choice, and building the store in your own cloud account, is what our cloud architecture and AWS engineering work is for.
Where the effort and the cost really go
Three decisions drive most of the long-term cost of a Knowledge Store, and all three are made early.
The embedding model. Every vector is produced by a particular model. Change the model and every vector has to be regenerated, which on a large store is an operational event with its own plan, budget and rollback. Choose deliberately, record which model produced each vector, and do not switch on a whim.
The number of relationship types. Every kind of link in the graph is a schema commitment and a future migration. A dozen well-chosen relationship types will serve most organisations; fifty is usually a sign that metadata has leaked into the graph. A link nobody queries is clutter.
What happens to old entries. Deleting is rarely right, because it destroys the audit trail. Superseded and retracted entries stay in the store and are filtered out of default reads. That costs a little storage and saves a great deal of explaining later.
Six questions to ask your team
If a Knowledge Store is being built or already runs in your organisation, these questions will tell you quickly how mature it is.
- Which tier is the source of truth?If the answer is "all of them", you have three sources of disagreement.
- When did we last rebuild the index from the records?Rehearse it every quarter. Find out it is broken in a drill, not in an incident.
- What happens when we change the embedding model?There should be a written plan: run old and new side by side, backfill, cut over, keep a way back.
- Can an agent read outside its scope?The store is an access-control system. Sample agent reads each month and check every one was allowed.
- Do we log reads as well as writes?Without read logs you cannot show what an agent knew when it made a decision.
- Whose keys encrypt it?Use customer-managed keys. The store holds what makes your organisation different from its competitors.
Common questions about AI agent memory
What is AI agent memory?
It is the information an AI agent can look up at the moment it answers: policies, past decisions and what the organisation has learned. Unlike what a model absorbed in training, it can be updated, scoped to the people allowed to see it, and audited. In this article the system that holds it is called the Knowledge Store.
Is a vector database enough for AI agent memory?
For a prototype, often yes. In production, a vector database on its own cannot tell which of several similar entries is current, who approved it, or whether this agent is allowed to read it. That is why the pattern adds a graph and a record store, with the record store as the single source of truth.
Which AWS services are used to build AI agent memory?
A common combination is Amazon DynamoDB or Aurora PostgreSQL for the records, Amazon S3 Vectors or OpenSearch Serverless for similarity search, and Amazon Neptune for relationships, with Amazon Bedrock Knowledge Bases as the retrieval layer the agents call.
Should we use S3 Vectors or OpenSearch Serverless?
S3 Vectors suits large stores where cost matters more than the last few milliseconds of latency. OpenSearch Serverless suits workloads that also need keyword and hybrid search at high query rates. Some teams use both, keeping the bulk of the store in S3 Vectors and the busiest content in OpenSearch.
Can the same design run on Azure or Google Cloud?
Yes. On Azure, Cosmos DB holds the records and supports vector search, Azure AI Search handles search-heavy workloads, and Cosmos DB for Apache Gremlin holds the graph. On Google Cloud, Spanner can hold all three tiers in one service: tables, vector search and Spanner Graph.
The short version
Agents need three kinds of memory: similarity, relationships and exact records. Make the records the single source of truth, derive the other two from them, and write in that order. Read wide, then narrow, then exact, and log every read. Let entries be superseded or retracted, never silently deleted. Then build it on whichever cloud your data already calls home.
Where the Knowledge Store sits in the wider architecture is the subject of Microservices Break Under Agents, and Why AIM Fits. The learning process that fills it is described in OKL and the Organizational AI Brain. If you are building on Amazon Bedrock, our LLM platform on AWS Bedrock and data and knowledge platforms pages describe how we deliver it, and our AI engineering practice builds the agents that read from it.