Memory plane
Retrieval is not memory.
The first week is in the org, reading what already exists. A vector store finds text that looks like your question. It does not know what was decided last quarter, or that this customer already refused the upsell twice. Three layers, kept separate. The honest no: if the work is a software factory with a named delivery date, we are the wrong partner.
Certified Partner since 2010 · MVP Hall of Fame · 200+ agents in production · UAE and US desks
What it is
Three layers, three jobs.
Conflating them is the most common architecture mistake we inherit. Each layer has a different owner, a different retention rule, and a different failure.
Retrieval
Hybrid, then reranked, then compressed.
Pure vector search fails on exact identifiers. Pure keyword search fails on meaning. Production retrieval runs both and then pays for precision at the end.
01
Decompose the question
Intent, metadata filters, and where useful a hypothetical answer to embed. HyDE-style expansion lifts recall on vague questions.
02
Retrieve two ways at once
Embeddings for meaning, lexical search for part numbers and clause references. Fuse the two ranked lists rather than picking a winner.
03
Traverse the graph
Entity and theme indexing, LightRAG-style, so a question that spans documents does not lose the link between them.
04
Rerank and compress
A cross-encoder scores the shortlist properly, then weak passages are dropped. Less context, better answers, lower spend.
05
Check what you already know
Durable memory is read before the model answers, so the agent does not re-litigate a decision the business already made.
Mermaid source
flowchart TD
Q["Question"] --> D["Decompose: intent, filters, hypothetical answer"]
D --> V["Dense retrieval: embeddings"]
D --> K["Lexical retrieval: exact terms and identifiers"]
D --> G["Graph traversal: entities and themes"]
V --> P["Fused candidate pool"]
K --> P
G --> P
P --> R["Cross-encoder rerank"]
R --> C["Compress to the passages that earn their tokens"]
C --> M["Durable memory: what we already decided"]
M --> A["Answer or tool call"]Why it matters
Cold-start agents get fired by their users.
Your team will forgive a slow answer. They will not forgive being asked the same question every morning.
In an org
A hospital network or a bank.
Same three layers, different words for them.
How Mindcat helps
Where we would start in your org.
Context before agency. Dirty data stalls agents, and no framework fixes that.
Questions
What we actually say.
- Can a bigger context window replace this?
- No. A long prompt is working context. It expires with the session and it is not governed, searchable, or auditable. Large windows are wonderful for reading a whole repository or contract set in one pass, which is a different job.
- Do we need Data Cloud before an agent?
- You need grounded, permissioned context. Sometimes that is Data Cloud. Sometimes it is clean knowledge and a sharing model that already works. We check before quoting a build.
- Graph or vectors?
- Both, for different questions. Relationships in a graph, similarity in a vector store, exact identifiers in lexical search. Anyone selling one of the three as the whole answer has not shipped this.
- How long do you keep agent memory?
- As long as the purpose justifies, with deletion defined up front. We treat a memory record like a CRM field, not like a log file.
Next
What to read next.
The brief · one email
Name what the agent must never forget.
We will tell you which layer that belongs in, and what it may not store.
