PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 24, 20260 citationsOpen Access

Benchmark Recall versus Deployed Consistency: An Empirical Comparison of Memory Strategies for Local Small-Model NPCs

View Full Paper
QCQ. Chen

Key Points

  • This research aims to evaluate different memory management strategies for local small-model non-playable characters (NPCs).
  • Three memory strategies were compared: sliding-window recency, embedding-based retrieval, and LogMem.
  • The evaluation involved testing with 10 REALTALK conversations and one LoCoMo conversation under fixed token budgets.
  • Mean judge scores were recorded across various token limits: 1K, 2K, and 4K.
  • Embedding-based retrieval achieved the highest mean scores for benchmark recall across all budgets.
  • The study emphasizes a distinction between benchmark recall and NPC consistency, underscoring the need for enhanced evaluations in games.
  • Findings suggest that future evaluations should balance these two different targets.

Abstract

This preprint studies memory management for local small-model NPCs under fixed prompt budgets. We compare three representative strategies: sliding-window recency, embedding-based retrieval, and a query-independent compression method called LogMem. The evaluation uses 10 REALTALK conversations as the main testbed and one LoCoMo conversation as a sparse-dialogue control, under fixed total budgets of 1K, 2K, and 4K tokens per answer call. On benchmark-style factual recall, embedding-based retrieval achieves the highest mean judge score at every tested budget. At the same time, the paper argues that benchmark recall and deployed NPC consistency are different targets, and that future game-facing evaluations should measure both. The practical motivation for this work comes from the design and deployment of memory systems for persistent characters in the Veridia interactive narrative platform. Project website: https://veridia.games Files included:- PDF manuscript- LaTeX/source submission package

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Q. Chen (2026) studied this question.

synapsesocial.com/papers/69eb0a2e553a5433e34b4696https://doi.org/10.5281/zenodo.19695320
Ask AI
Helpful
Bookmark
Share
View Full Paper