Staff+ Software Engineer, Safeguards Human Review Tooling
Anthropic
San Francisco, CAOn-siteFull-time
About this role
<div class="content-intro"><h2><strong>About Anthropic</strong></h2> <p>Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.</p></div><h2>About the role</h2> <p>The Safeguards team is responsible for ensuring our models and products are developed and deployed safely. We&39;re looking for engineers for our Review Tooling team, which builds the systems that humans — and increasingly Claude — use to investigate potential harms and take enforcement actions across Anthropic&39;s first-party products and third-party cloud platforms.</p> <p>This is a foundational role: as one of the first engineers on this new team, you&39;ll own the tools our safety investigators rely on to understand what&39;s happening on our platforms and act on it, as well as the platform underneath those tools. That platform includes analytics capabilities, privacy-preserving primitives that keep review workflows compatible with our data retention commitments, and a sandbox environment where new review interfaces and workflows can be built and iterated quickly. As model capabilities and usage grow, you&39;ll also drive how we scale review through automation — building systems where Claude meaningfully extends what human reviewers can do, while keeping people in the loop where their judgment matters most.</p> <p>These are internal tools, but they are anything but low-stakes: the speed, clarity, and reliability of this tooling directly determines how quickly Anthropic can identify harmful behavior, make sound enforcement decisions, and feed signal back into model training. You&39;ll partner closely with policy, operations, data scienc