Why my agent returns zero memories for new vendors

# agents# ai# software
Why my agent returns zero memories for new vendorspavani praharshitha

The most important line of code in my accounts payable agent is the one that makes it forget...

The most important line of code in my accounts payable agent is the one that makes it forget everything. When an invoice arrives from a vendor we have never paid, the agent skips memory recall and returns an empty list. That looks like a missing feature. It is the feature.

**

What LedgerMind does

**
LedgerMind is an accounts payable decision agent. An invoice comes in, and the agent approves it, flags an exception, holds the payment, or sends it to a human.

The pipeline is short:

Invoice → Purchase Order → Goods Receipt → Memory Recall → Risk Evaluation → Decision → Memory Retain

The first three steps are ordinary AP controls: find the PO, find the goods receipt, run a three-way match. The last four are where the interesting decisions live.

Every run can execute with memory or without, so I can see, invoice by invoice, where memory changed the outcome.

Memory comes from Hindsight, an open-source agent memory system, via its TypeScript client. The agent makes two calls: recall before deciding and retain after. This post is about how carefully I had to constrain both.

Where Hindsight sits in the stack
The Express backend calls it for recall and retain.

**

An empty recall is a valid answer

**
Most writing about agent memory focuses on retrieval quality: more context, ranked well, smarter agent. In AP that framing is dangerous. Payment decisions are trust decisions, and trust comes from a specific vendor's specific history.

A memory layer that helpfully returns something for a vendor with no history is manufacturing trust that doesn't exist.

Semantic search always returns nearest neighbors. Ask a shared bank about a vendor it has never seen and you get memories about other vendors with similar categories, terms, and amounts. If downstream logic treats those as evidence, a brand-new supplier inherits the reputation of its lookalikes, which is exactly what a bank-detail fraud attempt wants.

So recallVendorMemory refuses to ask:

The vendor master says whether we have ever paid this supplier. If not, there is nothing to recall, and the response carries its own source value so the trace and audit log can tell "no history" apart from "Hindsight failed."

In payment, vendor lands in evaluate, a zero HUMAN REVIEW at 61% confidence. New vendors are gated by finance staff, which is the right default.

Recall that can't wander
Even for known vendors I don't trust raw recall output. The bank is shared by every vendor and also contains the agent's own past decisions. That caused two problems: results about other vendors, and the agent citing itself, where a recall for today's invoice surfaces the decision it made on an earlier run. That is circular evidence.

The query asks Hindsight to ignore current decisions, but I don't rely on the query alone. Results are filtered in code:
const rawMemories = (result.results || [])
.map(r => ({
text: r.text || r.content || '',
type: r.type || 'memory'
}))
.filter(m =>
m.text &&
m.text.toLowerCase().includes(vendorNameLower)
)
.filter(m => !m.text.includes(invoice.invoiceNumber));

const memories = rawMemories.slice(0, 4);
A memory must mention the vendor by name, must not mention the current invoice number, and only the top four survive.

The string matching is blunt, which is why vendor identity comes from the vendor master rather than from whatever the invoice says. But blunt and auditable beats clever and opaque here. When someone asks why the agent held a payment, I want the answer to be four snippets I can read.

What memory buys: a baseline for surprise
The clearest case is a bank-account change. The bank holds a memory that Ravi Steel & Components completed 14 payments to the account ending 4821 with no beneficiary changes. A new invoice arrives for ₹9,58,632 with bank details ending 7719.

Without memory, the agent sees a changed account and routes to human review at 72% confidence. With memory, it knows what normal looks like for this vendor:
const priorPayments =
useMemory
? (rememberedPaymentCount || invoice.priorPaymentCount || 0)
: 0;

const firstBankChange =
bankChanged && priorPayments >= 10;
Fourteen clean payments to one account, then a different one, is a far stronger signal than a changed account on a vendor with three payments. The agent raises a FIRST_BANK flag, drops vendor trust from 90 to 30, and sets PAYMENT BLOCKED at 99% confidence.

The recommendation names the specifics: first change after 14 successful payments to ••••4821, verify the new account using the contact already on file.

Note "already on file." The verification contact comes from history, not from the invoice, because the invoice is the thing being challenged.

The payment count is the part I like least. It is parsed from recalled text with a regex, which is as fragile as it sounds. It's the first thing I'd move to structured metadata attached at retain time. Recalled prose is good for explaining a decision. For a threshold check, use a number.

The second behavior is a recurring exception. ABC Technologies once invoiced 22 monitors against 20 received, and AP resolved it by requesting a corrected invoice. When the same mismatch reappears, the agent recalls that resolution and recommends a correction instead of escalating.

Writing memory: retain surprises, not everything
If the agent retained every decision, routine approvals would crowd out the memories that carry signal. So it writes back only when something meaningful happened:
const meaningfulLearning =
useMemory &&
['PAYMENT BLOCKED', 'EXCEPTION', 'HUMAN REVIEW']
.includes(decision.status);

const alreadyRetained = store.audit.some(
a =>
a.action === 'AGENT_MEMORY_RETAINED' &&
a.invoiceId === invoice.id
);

if (meaningfulLearning && !alreadyRetained) {
retain = await retainEvent(
/* vendor, invoice, decision, evidence */
);
}
Clean approvals are not written back. Blocks, exceptions, and human-review routes are, with the evidence that triggered them and metadata for filtering later.

The alreadyRetained check exists because people click "Run" more than once. Without idempotence, one blocked invoice produced several near-identical memories that skewed recall toward it. This is also why recall excludes the current invoice number: decision events and vendor history share one bank, and the agent has to keep them apart.

Lessons
Treat empty recall as a first-class result. Memory systems return nearest neighbors by design. Decide up front which cases require nothing, and enforce it before the query, not after.

Filter recall in code, not just in the prompt. The prompt helps. Checking the results is what makes behavior reliable and reviewable.

Keep the evidence set small enough to read. Four memories are enough to justify a decision and few enough for an auditor to check.

Be selective about writes. Tomorrow's recall quality depends on what you retain today. Store events that changed a decision.

Use structured fields for thresholds. Prose memory is for reasoning. Anything a rule compares against a number should be stored as one.

If you're evaluating memory for agents that make consequential decisions, the Hindsight GitHub repository is a good place to start Hindsight GitHub repository, and Vectorize's overview of agent memory vs RAG explains the underlying idea.
https://vectorize.io/articles/agent-memory-vs-rag

In my experience the hard part isn't getting an agent to remember. It's deciding what it should be allowed to remember, and being comfortable when the honest answer is nothing.