Research Engineer, Interpretability
Anthropic
San Francisco, CAOn-siteFull-time
About this role
<div class="content-intro"><h2><strong>About Anthropic</strong></h2> <p>Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.</p></div><h2>About the role:</h2> <p>When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?"</p> <p>The Interpretability team at Anthropic is working to reverse-engineer how trained models work because we believe that a mechanistic understanding is the most robust way to make advanced systems safe.</p> <p>Think of us as doing "neuroscience" of neural networks using "microscopes" we build - or reverse-engineering neural networks like binary programs.</p> <p>More resources to learn about our work:&nbsp;</p> <ul> <li><a href="https://transformer-circuits.pub/">Our research blog</a> - covering advances including <a href="https://transformer-circuits.pub/2024/scaling-monosemanticity/">Monosemantic Features</a> and <a href="https://transformer-circuits.pub/2025/attribution-graphs/methods.html">Circuits</a></li> <li><a href="https://www.youtube.com/watch?v=TxhhMTOTMDg">An Introduction to Interpretability</a> from our research lead, <a href="https://colah.github.io/about.html">Chris Olah</a></li> <li><a href="https://www.darioamodei.com/post/the-urgency-of-interpretability">The Urgency of Interpretability</a> from CEO Dario Amodei</li> <li><a href="https://www.anthropic.com/research/engin