[Journal Club] Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations

Abstract

大規模言語モデル(LLM)の内部状態は、高次元の活性化ベクトルとして表現されており、モデルの推論・判断・潜在的な認識に関する豊かな情報を含んでいます。本資料では、この不透明な内部表現を自然言語で読み解く新たな解釈可能性手法 Natural Language Autoencoders(NLA)について、Anthropicの研究論文「Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations(Fraser-Taliente, Kantamneni, Ong et al. - 2026)」に基づき、手法・評価・監査応用・限界を体系的に解説します。

  • 📝:https://transformer-circuits.pub/2026/nla/
  • 🐙:https://github.com/kitft/natural_language_autoencoders
  • 🌐:https://www.neuronpedia.org/nla

Date
Jun 5, 2026 12:00 PM — 12:00 PM
Event
Speaker Deck
Location
Online

Speaker Deck

大規模言語モデル(LLM)の内部状態は、高次元の活性化ベクトルとして表現されており、モデルの推論・判断・潜在的な認識に関する豊かな情報を含んでいます。本資料では、この不透明な内部表現を自然言語で読み解く新たな解釈可能性手法 Natural Language Autoencoders(NLA)について、Anthropicの研究論文「Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations(Fraser-Taliente, Kantamneni, Ong et al. - 2026)」に基づき、手法・評価・監査応用・限界を体系的に解説します。

  • 📝:https://transformer-circuits.pub/2026/nla/
  • 🐙:https://github.com/kitft/natural_language_autoencoders
  • 🌐:https://www.neuronpedia.org/nla

資料 - Slides

Shunsuke Kitada, Ph.D.
Shunsuke Kitada, Ph.D.
Research Scientist working on Vision & Language with Deep Learning

My research interests include deep learning-based natural language processing, computer vision, medical image processing, and computational advertising.