Speaker
Description
Artificial intelligence systems can do more and more intricate tasks, but still, the inside stuff behind why they behave the way they do can be hard to get a grip on. And really, if humans built AI, how could we not know the way it works ? The basic point is the difference between making a learning system, versus directly programming the internal computations it ends up forming.
In this talk we’re looking at Mechanistic Interpretability (MI)—a kinda emerging method that tries to reverse-engineer neural networks by pushing their behavior back toward internal representations, features, circuits, and computations. Along the way there’s a neat tie-in to neuroscience too: instead of treating intelligence as something we only infer from behavior, could we understand it by finding the hidden mechanisms that actually generate it ?
Then we’ll shift to why it matters for AI safety and security. Like, if we had a clearer view of model internals, would that give us a better perspective for examining hidden capabilities, deceptive conduct, instruction conflicts, or other weird behaviors that might not show up just from the outputs alone
Finally, we’ll cover the current limits of MI, and why understanding more and more capable AI systems may end up being one of those defining scientific , and security challenges for the future, even if that sounds a bit far off right now.
Session author's bio
I’m a B.Tech. Information Technology UG student at IIIT Allahabad, working at the junction of cybersecurity and artificial intelligence. Earlier, I worked as an Undergraduate Researcher at C3iHub, IIT Kanpur, on IoT and cybersecurity-related research, and right now I’m associated with an Indo-Swiss research initiative with ETH Zurich looking at AI for sustainable and resilient supply chains.
My broader research interests include AI and LLM security, trustworthy AI, and mechanistic interpretability, though honestly, I’m particularly interested in what’s really going on under the hood. More specifically, I want to understand the internal mechanisms of AI systems and explore how that kind of insight can contribute to better security and accountability as AI systems become increasingly capable.
| In Person Attendance | In-person |
|---|---|
| Agree to Privacy Policy and Notice | I agree |
| Please confirm that there are included headshots of all speakers in their profiles | Yes |
| Level of Difficulty | Beginner |