Discussion
Forum
The SAIRC Discussion Forum is a space for AI enthusiasts to share what they're thinking about. No formal research paper required. Posts can be submitted anonymously and span a wide range of formats:
- Tutorials or deep-dives on AI topics
- Fresh perspectives: reframings or new ways of looking at something
- Novel research results communicated in plain language
- Resources and opportunities in AI (summer programs, tools, datasets)
- Thought experiments and speculative ideas
- Notes or study guides from courses
Recent Posts
The Underwhelming Universal Approximation Theorem
A research advisor brought up the Universal Approximation Theorem during a meeting once as a fun thought experiment. It stuck with me because I initially found it super interesting…
Mechanistic Interpretability and the Curses of Scaled Networks
Mechanistic interpretability is the practice of reverse-engineering deep-learning systems to understand their inner algorithms. If we can figure out what a neural network is actual…
What Students Want Teachers to Know About AI
Reposted with permission from original author(s).
Back in December I sat down with a handful of HS students on a couple different occasions to talk about AI — their thoughts, attit…
How Does Persistent Memory Affect AI Safety?
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.
Written with Maksym Andriushchenko, th…
In an AI World, What's the Work?
Reposted with permission from original author(s).
More than a decade ago, immersed in research on competency-based education, I read Leaders of Their Own Learning, a comprehensive …
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.
A security evaluation of the emerging …
AI Shouldn't Be Doing More Philosophy Than Students
Reposted with Permission from Mike Taubman and AI Waypoints.
Last weekend I found out that AIs can now reincarnate. This got me thinking about both the high school juniors I teach …
MLSN #19: Honesty, Disempowerment, & Cybersecurity
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.
An issue of the ML Safety Newsletter, …
Labor Market Impacts of AI: A New Measure and Early Evidence
Full credit goes to the original author, linked below. All blog posts were reposted either with permission of the author, or by anonymous submission by SAIRC members like yourself.…
HalluHard: A Hard Multi-Turn Hallucination Benchmark
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author.
Built with Maksym Andriushchenko, Hall…
Interpretability Research Already Has a Framework for Actionability
This post was either an anonymous submission of an interesting paper or was written by a student; full credit remains with the author (linked).
Prompted by a Chris Olah talk years…
There's No Token for the Way the End of High School Feels
Reposted with permission by Mike Taubman.
Today in our AI literacy class, Scott Kern and I helped students open the hood to see how AI works, one pillar of the AI Driver's License …
Painless Activation Steering (PAS): Automated, Lightweight Post-Training for LLM Behavior
Reproduced with permission from Sasha Cui
We're releasing "Painless Activation Steering (PAS)," a fully automated approach to steer large language models after training—without mo…
(When) Is Mechanistic Interpretability Identifiable?
I recently finished a paper, "Characterizing Mechanistic Uniqueness and Identifiability Through Circuit Analysis," alongside a group of three others and a mentor. This post discuss…
Mamba's Memory Problem
Full credit goes to the original author, linked below. All blog posts were reposted either with permission of the author, or by anonymous submission by SAIRC members like yourself.…