Neel Nanda
Neel Nanda is a prominent researcher in the field of AI Safety and mechanistic-interpretability. He serves as the Mechanistic Interpretability Team Lead at google-deepmind.
Key Contributions & Work
- Leads the Mechanistic Interpretability Team at google-deepmind, focusing on understanding the internal mechanisms of large language models.
- Creator of TransformerLens, a popular library for mechanistic interpretability research.
- Active contributor to the open-source interpretability community, promoting transparency and safety in AI development.
Media & Interviews
- Featured in the Google DeepMind podcast episode discussing the inner workings of AI models and the importance of interpretability for safety.
- AI Interpretability: Unpacking Black Box Models for Safety and Science