Speaking twice at Dreamforce · Sept 15-17 →

Memory plane

Retrieval is not memory.

The first week is in the org, reading what already exists. A vector store finds text that looks like your question. It does not know what was decided last quarter, or that this customer already refused the upsell twice. Three layers, kept separate. The honest no: if the work is a software factory with a named delivery date, we are the wrong partner.

Get the written assessment

Certified Partner since 2010 · MVP Hall of Fame · 200+ agents in production · UAE and US desks

Data Cloud

What it is

Three layers, three jobs.

Conflating them is the most common architecture mistake we inherit. Each layer has a different owner, a different retention rule, and a different failure.

Retrieval

Hybrid, then reranked, then compressed.

Pure vector search fails on exact identifiers. Pure keyword search fails on meaning. Production retrieval runs both and then pays for precision at the end.

  1. 01

    Decompose the question

    Intent, metadata filters, and where useful a hypothetical answer to embed. HyDE-style expansion lifts recall on vague questions.

  2. 02

    Retrieve two ways at once

    Embeddings for meaning, lexical search for part numbers and clause references. Fuse the two ranked lists rather than picking a winner.

  3. 03

    Traverse the graph

    Entity and theme indexing, LightRAG-style, so a question that spans documents does not lose the link between them.

  4. 04

    Rerank and compress

    A cross-encoder scores the shortlist properly, then weak passages are dropped. Less context, better answers, lower spend.

  5. 05

    Check what you already know

    Durable memory is read before the model answers, so the agent does not re-litigate a decision the business already made.

Production retrieval pipeline. Decompose the question, retrieve densely and lexically in parallel, fuse the rankings, rerank with a cross-encoder, compress to what the model actually needs, then check durable memory before answering.
Mermaid source
flowchart TD
  Q["Question"] --> D["Decompose: intent, filters, hypothetical answer"]
  D --> V["Dense retrieval: embeddings"]
  D --> K["Lexical retrieval: exact terms and identifiers"]
  D --> G["Graph traversal: entities and themes"]
  V --> P["Fused candidate pool"]
  K --> P
  G --> P
  P --> R["Cross-encoder rerank"]
  R --> C["Compress to the passages that earn their tokens"]
  C --> M["Durable memory: what we already decided"]
  M --> A["Answer or tool call"]

Why it matters

Cold-start agents get fired by their users.

Your team will forgive a slow answer. They will not forgive being asked the same question every morning.

In an org

A hospital network or a bank.

Same three layers, different words for them.

Questions

What we actually say.

Can a bigger context window replace this?
No. A long prompt is working context. It expires with the session and it is not governed, searchable, or auditable. Large windows are wonderful for reading a whole repository or contract set in one pass, which is a different job.
Do we need Data Cloud before an agent?
You need grounded, permissioned context. Sometimes that is Data Cloud. Sometimes it is clean knowledge and a sharing model that already works. We check before quoting a build.
Graph or vectors?
Both, for different questions. Relationships in a graph, similarity in a vector store, exact identifiers in lexical search. Anyone selling one of the three as the whole answer has not shipped this.
How long do you keep agent memory?
As long as the purpose justifies, with deletion defined up front. We treat a memory record like a CRM field, not like a log file.

The brief · one email

Name what the agent must never forget.

We will tell you which layer that belongs in, and what it may not store.