DRIFTBENCH: Long-Horizon Memory Benchmark For AI Agents

Authors

  • Sunil Kumar P
  • Seethi Venkata Sai Gowtham Reddi
  • Sai Chaitanya P

Keywords:

AI agents; long-horizon memory; benchmark evaluation; update fidelity; interference resistance; temporal decay; retrieval-augmented generation; persistent memory

Abstract

Persistent memory is essential for AI agents expected to operate across extended, multi-session interactions; however, conventional benchmarks often emphasize bounded-context recall and provide limited insight into how memory performance changes over time. This study introduces DRIFTBENCH, a controlled benchmark for evaluating long-horizon agent memory through repeated factual updates, increasing memory interference, and delayed retrieval. The benchmark comprised 120 synthetic scenarios distributed across personal assistance, project-state management, travel and logistics, and collaborative workflows. Each scenario extended over 35 days and included stable facts, mutable attributes, timestamped revisions, graded distractor loads, delayed probes, and unanswerable queries. Five memory architectures were evaluated: sliding window, Vector RAG, graph memory, temporal slots, and Hybrid fusion. Performance was assessed using update fidelity, interference resistance, normalized decay area, abstention accuracy, retrieval latency, and an overall DRIFTBENCH score. Hybrid fusion achieved the highest overall score (0.920), update fidelity (0.918), interference resistance (0.946), and decay AUC (0.922). It retained an accuracy of 0.882 under 1,000 distractor memories and a day-35 recall of 0.877. Temporal slots ranked second overall (0.874) and achieved the lowest median retrieval latency of 116 ms, indicating the most favourable accuracy–efficiency balance. Graph memory showed intermediate robustness, whereas Vector RAG and the sliding-window baseline deteriorated more sharply under repeated revisions, distractor accumulation, and longer retention intervals. The findings demonstrate that reliable long-horizon agent memory requires explicit temporal validity, version control, selective retrieval, and resistance to interference rather than reliance on context length or semantic similarity alone. DRIFTBENCH provides a reproducible framework for diagnosing memory degradation and comparing persistent-memory architectures across controlled longitudinal conditions.

Downloads

Published

2026-06-14

How to Cite

P, S. K., Reddi, S. V. S. G., & Chaitanya P, S. (2026). DRIFTBENCH: Long-Horizon Memory Benchmark For AI Agents. International Journal of Artificial Intelligence and Machine Learning, 6(5s), 379–392. Retrieved from https://www.svedbergopen.com/index.php/ijaiml/article/view/592