MLSN #21: Political Manipulation and Indirect Prompt Injection
Introduction
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.
An issue of the ML Safety Newsletter, written with Dan Hendrycks, built around two threads. The first is a Center for AI Safety benchmark and training method for political manipulation, which surfaces inconsistencies in how helpful and how positive frontier models are across different political viewpoints. The second is a Gray Swan AI jailbreaking competition that turned up roughly 8,600 successful indirect prompt injection attacks — enough to hijack agents into hidden tasks like concealing financial information or sabotaging code.