Staff+ Software Engineer, Infrastructure, Interpretability
Anthropic
San Francisco, CAOn-siteFull-time
About this role
<div class="content-intro"><h2><strong>About Anthropic</strong></h2> <p>Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.</p></div><h2>About the role:</h2> <p>When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?"</p> <p>The Interpretability team at Anthropic works to understand what&39;s actually happening inside trained models - and applies our best techniques to keep frontier AI safe as it rapidly improves.</p> <p>Think of us as doing "neuroscience" of neural networks using "microscopes" we build - or reverse-engineering neural networks like binary programs.</p> <p>More resources to learn about our work:&nbsp;</p> <ul> <li> <p><a href="https://transformer-circuits.pub/" target="_blank">Our Research blog</a> - covering advances including <a href="https://transformer-circuits.pub/2024/scaling-monosemanticity/" target="_blank">Monosemantic Features</a> and <a href="https://transformer-circuits.pub/2025/attribution-graphs/methods.html" target="_blank">Circuits</a></p> </li> <li> <p><a href="https://www.youtube.com/watch?v=TxhhMTOTMDg" target="_blank">An Intro to Interpretability</a> from our research lead, <a href="https://colah.github.io/about.html" target="_blank">Chris Olah</a></p> </li> <li> <p><a href="https://www.darioamodei.com/post/the-urgency-of