MIT, Harvard & Northeastern U's Sparse Probing Aims at 'Finding Neurons in a Haystack'
One of the most vexing public concerns regarding powerful AI systems such as large language models (LLMs) is their lack of interpretability: How can people trust a model’s output when they cannot understand how it was formed? In the new paper Finding Neurons in a Haystack: Case Studies with Sparse Probing, a research team from
Source: syncedreview.com