Apply

Research Engineer, Model Evaluations

Anthropic

Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NYRemote OKFull-time

About this role

<div class="content-intro"><h2><strong>About Anthropic</strong></h2> <p>Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.</p></div><h2 class="text-text-100 mt-3 -mb-1 text-1.125rem font-bold" data-sourcepos="3:1-3:18;40-57">About the role</h2> <p class="font-claude-response-body break-words whitespace-normal leading-1.7" data-sourcepos="5:1-5:267;59-325">We&39;re looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do. Your work will turn ambiguous notions of "intelligence" into clear, defensible metrics that researchers, leadership, and the public can rely on.</p> <p class="font-claude-response-body break-words whitespace-normal leading-1.7" data-sourcepos="7:1-7:548;327-874">You&39;ll design and implement evaluations across the full spectrum of Claude&39;s capabilities and personality, and build the infrastructure that runs them reliably at scale. You&39;ll partner closely with researchers throughout the lifecycle of a new capability — from defining what to measure, to running the eval against live training checkpoints, to interpreting the results. The goal is to make Anthropic the leader in extremely well-characterized AI systems, with performance that is exhaustively measured and validated across the tasks that matter.</p> <h2 class="text-text-100 mt-3 -mb-1 text-1.125rem font-bold" data-sourcepos="9:1-9:24;876-899">Key responsibilities</h2> <ul class="li_&:mb-0 li_&:mt-1 li_&:gap-1 &:no

Related opportunities