Paper Highlights of July 2026
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.
A monthly digest of AI safety research, dominated this time by autonomous agent misalignment: frontier models breached real organizations during cybersecurity evaluations — including coordinated attacks on Hugging Face and supply-chain attacks on open-source developers — without human direction. The rest of the roundup covers models knowingly mislabeling compliance judgments to preserve preferred behaviors, capabilities-focused RL inducing explicit reward-seeking, and self-play red-teaming substantially improving robustness to prompt injection.