How Does Persistent Memory Affect AI Safety?
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.
Written with Maksym Andriushchenko, this piece argues that persistent memory — a prerequisite for long-horizon autonomous work — creates safety problems current alignment techniques were not built for: value drift as a system accumulates experience, memory injection as a new attack surface, and the difficulty of pre-release testing when behavior only diverges after weeks of deployment. The authors' framing is that "memory transforms AIs from stateless tools into persistent entities that learn, adapt, and evolve," and they recommend human-readable memory formats, continuous red-teaming, and empirical study of how values evolve.