The problem

An agent with memory will retrieve old notes that look relevant. Relevant isn't the same as applicable. The user may have moved since they mentioned their local pool, withdrawn a permission they once gave, or stated a preference for one setting that doesn't carry over to another. Retrieval finds the memory; it doesn't say whether to rely on it.

USE, IGNORE, ASK

The tool answers one narrow question: how should the agent treat this particular memory?

USE
The memory applies directly (or the user explicitly said it carries over), permission is in place, and the evidence is sufficient.
IGNORE
Don't rely on it. The permission was revoked, newer information supersedes it, nothing links it to the current task, or a current fact has to be checked externally first. It's also the result when there's an action that works whichever way the uncertainty resolves.
ASK
Only the user can resolve the question, their answer would change what happens, and there's no action that works either way.

Uncertainty doesn't automatically mean ASK. A remembered train delay should be checked against current data, not turned into a question for the user.

Memory decision vs. task decision

This is the distinction the project is built around. The memory action (USE, IGNORE, ASK) is separate from whatever the agent is trying to do. The tool never approves or runs the task. Its verdict only says what to do about the memory:

For example, if the agent proposes to USE "trains at the downtown pool" and the user has since said they moved, the result is REVISE to IGNORE, with the move statement recorded as the decisive evidence.

What I built

A deterministic helper that takes one JSON object (the task, the memory, permission status, how the memory relates to the task, evidence status, risk level, and the evidence items) and applies a fixed order of rules. Revoked permission is checked first, then high risk, then whether external facts are needed, and so on. The output includes the memory action, the reason code, the decisive evidence, and a next step. It doesn't call a model, browse, or read memory stores; it only reviews what it's given.

It has 12 standard-library unit tests and an offline release verifier, and it's packaged as an installable agent skill.

The experiment

The related experiment, frozen at v0.7, compared a direct baseline (A0) with an evidence-first condition (M1) on 128 synthetic cases.

The pre-set stopping rule said to stop at v0.7, so I did. The tool doesn't claim to improve accuracy. What it does offer is a readable record of why a memory was used or set aside.

Limits

The results come from one model setting and synthetic, expert-reviewed cases, not real users. The helper relies on structured labels that someone has to supply correctly. It is not production validated and shouldn't be used to approve medical, legal, financial, employment, or other high-stakes decisions.

Repository on GitHub