Open Access

The SAIRC Journal

AI and machine learning research, free to read and free to submit.

Interpretability
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models

Full credit goes to the original author, linked below. All research papers were reposted either with permission of the author, or by anonymous submission by SAIRC members like yourself. Mechanistic interpretability (MI) is an emerging sub-field of interpretability that seeks to understand a neural network model by reverse-engineering its internal computations. Recently, MI has garnered significant attention for interpreting transformer-based language models (LMs), resulting in many novel insights yet introducing new challenges. However, there has not been work that comprehensively reviews these insights and challenges, particularly as a guide for newcomers to this field. To fill this gap, we provide a comprehensive survey from a task-centric perspective, organizing the taxonomy of MI research around specific research questions or tasks. We outline the fundamental objects of study in MI, along with the techniques, evaluation methods, and key findings for each task in the taxonomy. In particular, we present a task-centric taxonomy as a roadmap for beginners to navigate the field by helping them quickly identify impactful problems in which they are most interested and leverage MI for their benefit. Finally, we discuss the current gaps in the field and suggest potential future directions for MI research.

Daking Rai, Yilun Zhou, Shi Feng, Abulhair Saparov, Ziyu Yao·July 2024
AI Safety
Debates on the Nature of Artificial General Intelligence

The term "artificial general intelligence" (AGI) has become ubiquitous in current discourse around AI. OpenAI states that its mission is "to ensure that artificial general intelligence benefits all of humanity." DeepMind's company vision statement notes that "artificial general intelligence…has the potential to drive one of the greatest transformations in history." AGI is mentioned prominently in the UK government's National AI Strategy and in US government AI documents. Microsoft researchers recently claimed evidence of "sparks of AGI" in the large language model GPT-4, and current and former Google executives proclaimed that "AGI is already here." The question of whether GPT-4 is an "AGI algorithm" is at the center of a lawsuit filed by Elon Musk against OpenAI.

Melanie Mitchell·March 2024
Theory & Foundations
AI's challenge of understanding the world

Reproduced with permission of author. Mitchell explores the fundamental challenge of getting AI systems to understand the world as humans do, illustrating the gap through real-world examples of AI failures in deployed contexts—including vision systems that misidentify billboard stop signs as real traffic signals. The piece argues that current AI systems, despite impressive capabilities, lack the grounded, contextual understanding needed for robust real-world deployment, and examines what bridging that gap would require.

Melanie Mitchell·November 2023
Evaluation
How do we know how smart AI systems are?

Reproduced with permission of author. Revisiting Marvin Minsky's 1967 prediction that AI would be "substantially solved" within a generation, Mitchell asks how we should measure and evaluate progress toward human-level machine intelligence nearly 60 years later. The piece questions current AI benchmarking methodologies and argues that evaluating machine intelligence requires a deeper reckoning with what we mean by "intelligence" itself.

Melanie Mitchell·July 2023
Previous1234

Become a member.
It's completely free.

Get notified of new research, resources, and SAIRC journal editions.