Aug 28 – 30, 2026
Asia/Kolkata timezone

Mechanistic Interpretability: Inside the Mind of AI

Aug 29, 2026, 5:55 PM
10m
Room 2

Room 2

Lightning Talk (10 min) Artificial Intelligence, Machine Learning, Data Science

Speaker

Pranav Mamane Mamane
iit2025021@iiita.ac.in

Description

Artificial intelligence systems can do more and more intricate tasks, but still, the inside stuff behind why they behave the way they do can be hard to get a grip on. And really, if humans built AI, how could we not know the way it works ? The basic point is the difference between making a learning system, versus directly programming the internal computations it ends up forming.

In this talk we’re looking at Mechanistic Interpretability (MI)—a kinda emerging method that tries to reverse-engineer neural networks by pushing their behavior back toward internal representations, features, circuits, and computations. Along the way there’s a neat tie-in to neuroscience too: instead of treating intelligence as something we only infer from behavior, could we understand it by finding the hidden mechanisms that actually generate it ?

Then we’ll shift to why it matters for AI safety and security. Like, if we had a clearer view of model internals, would that give us a better perspective for examining hidden capabilities, deceptive conduct, instruction conflicts, or other weird behaviors that might not show up just from the outputs alone

Finally, we’ll cover the current limits of MI, and why understanding more and more capable AI systems may end up being one of those defining scientific , and security challenges for the future, even if that sounds a bit far off right now.

Session author's bio

I’m a B.Tech. Information Technology UG student at IIIT Allahabad, working at the junction of cybersecurity and artificial intelligence. Earlier, I worked as an Undergraduate Researcher at C3iHub, IIT Kanpur, on IoT and cybersecurity-related research, and right now I’m associated with an Indo-Swiss research initiative with ETH Zurich looking at AI for sustainable and resilient supply chains.

My broader research interests include AI and LLM security, trustworthy AI, and mechanistic interpretability, though honestly, I’m particularly interested in what’s really going on under the hood. More specifically, I want to understand the internal mechanisms of AI systems and explore how that kind of insight can contribute to better security and accountability as AI systems become increasingly capable.

In Person Attendance In-person
Agree to Privacy Policy and Notice I agree
Please confirm that there are included headshots of all speakers in their profiles Yes
Level of Difficulty Beginner

Presentation materials

There are no materials yet.