Session
Did That Memory Actually Help? (Counterfactual Evals for Agent Memory)
Give a financial agent memory and it gets more personal. Whether it gets more correct is a separate question that almost nobody tests.
The same stored fact has different value depending on the decision. A user's emergency-fund target is essential context for one question and pure distraction for a debt-repayment question. Semantic relevance to the user does not imply relevance to the decision, and retrieval systems optimize for the former.
I will walk a financial-agent prototype through the same memory set across budgeting, debt repayment, and emergency-fund decisions, comparing semantic retrieval, user-state retrieval, full-context prompting, and decision-aware retrieval. Then the workflow: freshness checks, counterfactual removal to test whether one specific memory improved the answer, deterministic calculation tools instead of trusting the model with arithmetic, and transparent memory-use logs.
For anyone deploying memory in a regulated context, the audit trail is not a nice-to-have. If you cannot say which stored facts produced an answer, you cannot defend the answer.
Key Takeaways:
- A memory metadata schema and retrieval policies for high-stakes decisions
- Counterfactual removal testing: did this specific memory improve the decision?
- Measuring decision quality against token cost, latency, and privacy exposure
- Audit trails that explain which memories drove a recommendation
Ishween Kaur
Senior Engineer, Crypto and AI @SoFi
Santa Clara, California, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top