Your work will focus on recovering human-meaningful structure from learned representations, including sparse decompositions of activations, concept discovery, mechanistic analysis, and causal intervention, and on the interactive interfaces and evaluation methodology that let experts interrogate that structure and the given explanations. Demonstrated research productivity, as documented by publications, reports, presentations, and/or open-source software in relevant venues (NeurIPS, ICML, ICLR, CVPR, ACL, IEEE VIS, CHI, JMLR, etc.).