Apply

Full-Stack Software Engineer, Reinforcement Learning

Anthropic

San Francisco, CA | New York City, NYOn-siteFull-time

About this role

<div class="content-intro"><h2><strong>About Anthropic</strong></h2> <p>Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.</p></div><h2><strong>About the Role</strong></h2> <p>As a Full-Stack Software Engineer in RL, you&39;ll build the platforms, tools, and interfaces that power environment creation, data collection, and training observability. The quality of Claude&39;s next generation depends on the quality of the data we train it on — and the systems you build are what make that data possible.</p> <p>You&39;ll own product surfaces end-to-end — from backend services and APIs to the web UIs that researchers, external vendors, and thousands of data labelers use every day. You don&39;t need a background in ML research. What matters is that you can take an ambiguous, high-stakes problem and ship a polished, reliable product against it, fast.</p> <p>This team moves very quickly. Claude writes a lot of the code we commit, which means the bottleneck isn&39;t typing — it&39;s judgment, taste, and the ability to react to what researchers need next. You&39;ll iterate on data collection strategies to distill the knowledge of thousands of human experts around the world into our models, and you&39;ll do it in a loop that closes in hours and days, not quarters or months.</p> <p>Anthropic&39;s Reinforcement Learning organization leads the research and development that trains Claude to be capable, reliable, and safe. We&39;ve contributed to every Claude model, with significant impact on the autonomy and coding capabilities of our most advanced models. Our work spans teaching models

Related opportunities