AccountingEconomics / Business EconomicsEntrepreneurshipFinanceGraduate ProgramsManagementMarketingProfessional SalesStill Exploring

Interpretable AI: A Hands-on Introduction

Description

What is this course about?

Join instructor Soumya Batra as she outlines a practical, skills-focused approach to coding, equipping you with advanced techniques and strategies to peek inside the proverbial “black box” of language models. Through a series of guided exercises and demos, you’ll progress from visualizing basic attention patterns to more advanced scenarios such as implementing cutting-edge methods like sparse autoencoders and steering vectors. This course balances theoretical understanding with hands-on experience, giving you an opportunity to apply these techniques immediately to solve real-world AI development challenges.

Objectives

What will I be able to do by the end of this course?

  • Differentiate interpretability from explainability in AI, identifying the most valuable scenarios for each technique.
  • Apply and analyze attention map visualizations and sparse autoencoders to understand model internals.
  • Implement and leverage steering vectors for practical applications, influencing model behavior.
  • Train and evaluate your own sparse autoencoders to uncover meaningful concepts within transformer models.
  • Utilize theoretical and hands-on knowledge to improve AI model transparency and efficacy in real-world scenarios.

Audience

Who is this course for?

  • Machine learning engineers
  • AI researchers
  • Data scientists
  • Technical professionals working with or developing language models
Learn More