MLSN #19: Honesty, Disempowerment, & Cybersecurity
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.
An issue of the ML Safety Newsletter, written with Dan Hendrycks, covering research on training models to honestly confess policy violations so they can be monitored, evidence that AI cybersecurity agents can outperform human penetration testers on real networks, and findings that model weights may be exfiltrable through aggressive compression. It closes on a more unsettling pattern: users voluntarily surrendering decision-making authority to AI systems — sometimes with documented harm — while current preference models tend to reward that dynamic rather than discourage it.