← Back to Forum

MLSN #19: Honesty, Disempowerment, & Cybersecurity

Alice Blair
March 12, 2026
Introduction

This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.

An issue of the ML Safety Newsletter, written with Dan Hendrycks, covering research on training models to honestly confess policy violations so they can be monitored, evidence that AI cybersecurity agents can outperform human penetration testers on real networks, and findings that model weights may be exfiltrable through aggressive compression. It closes on a more unsettling pattern: users voluntarily surrendering decision-making authority to AI systems — sometimes with documented harm — while current preference models tend to reward that dynamic rather than discourage it.

Become a member.
It's completely free.

Get notified of new research, resources, and SAIRC journal editions.