Explainable AI Berlin’s Post

🔦 Past Paper Highlight: PRISM, interpreting polysemantic model features 📄 Capturing polysemanticity with PRISM: A multi-concept feature description framework 🔗 https://lnkd.in/dXXPthB4 💻 https://lnkd.in/dm-YR7Pp Most automated interpretability methods assume that individual neurons encode a single concept. The paper 'Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework' addresses this limitation by introducing a framework that generates multi-concept descriptions for LLM features. It introduces (i) a method for clustering high-activation inputs and labeling each cluster with an LLM to capture multiple concepts per neuron, (ii) a polysemanticity score to quantify how semantically diverse a neuron's associated concepts are, and (iii) a description score to evaluate how faithfully each concept label aligns with the neuron's activation distribution. Great work done by Laura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer, Anna Hedström, Marina Marie-Claire Höhne, Oliver Eberle

  • diagram

To view or add a comment, sign in

Explore content categories