The forum
We've featured 165 posts from members, read by roughly 2,800 readers per week. These posts consist of anonymously submitted works written by students, and works suggested by students which are written by other authors.
All posts
The Underwhelming Universal Approximation Theorem
A research advisor brought up the Universal Approximation Theorem during a meeting once as a fun thought experiment. It stuck with me because I initially found it super interesting…
Comparison of Convolutional & Feed-Forward Architectures on MNIST Digit Classification
I was unable to use Google Colaboratory for quite a bit, so it took me much longer than necessary to make this post. However, I'm finally able to log back in! This is the experimen…
The Convolutional Neural Network
The last few posts I've written about AI consciousness and infinite suffering have been fairly dire, so I decided to switch things up and write about something more practical: the …
AI Consciousness: A Biological Perspective
Most policy debates about AI revolve around its potential upsides: whether AI as an augmented decision-maker can solve existential risks like climate change or pandemics. But a dif…
Mechanistic Interpretability and the Curses of Scaled Networks
Mechanistic interpretability is the practice of reverse-engineering deep-learning systems to understand their inner algorithms. If we can figure out what a neural network is actual…
SCOTUS Simulator Using A Council of LLMs
I recently had the idea to evaluate an LLM council's stance on contested legal questions in a Supreme Court style. In this post, I'll feed the LLM Council various hypotheticals and…
Simulating War-Time Decisions with a Council of LLMs
'LLM Council' is a GitHub repo made by Andrej Karpathy which is intended to simulate a council of leading language models in a boardroom setting. It consists of council-member lang…
Autoregression & Next-Token Prediction
Every time a language model generates text, it's doing something surprisingly simple: predicting one token at a time, with each choice shaped by everything that came before. This p…
Would a Language Model Push You Off A Bridge?
In the context of this post, 'utilitarianism' is a consequentialist decision-making framework which operates under the idea that the best action produces the most pleasure for the …
Would a Language Model Push You Off A Bridge? Pt. 2
Utilitarianism, to recap, is a consequentialist decision-making framework which states that the best actions produce the most 'pleasure' for the greatest number of people. Deontolo…
(When) Is Mechanistic Interpretability Identifiable?
I recently finished a paper, "Characterizing Mechanistic Uniqueness and Identifiability Through Circuit Analysis," alongside a group of three others and a mentor. This post discuss…
What Every Student Should Know About AI Before They Graduate
In a few years, the students sitting in my classroom will be working alongside AI systems in nearly every industry: healthcare, finance, law, engineering, creative fields. Most of …
Teaching Against AI Sycophancy in the Research Writing Classroom
Republished with permission of author. Originally published at https://ruth.substack.com/ Two recent Science articles led me to rethink the impact of AI feedback in a research wri…
AI Safety Events & Training: 2026 Week 32 Update
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author. A weekly curated roundup of the AI saf…
Paper Highlights of July 2026
This post was submitted by a SAIRC member either as a recommended read or student-created post. All credit remains with the original author. A monthly digest of AI safety research…
Write something yourself
- Tutorials or deep-dives on an AI topic
- A reframing: a new way of looking at something people think is settled
- Research results written in plain language
- Resources and opportunities you found (summer programs, tools, datasets)
- Thought experiments and ideas you have not finished thinking through
- Notes or study guides from a course you took
Send it through the form and we will put it up, usually within a few days. If you would rather just email it, sairc.support@gmail.com works too.