Apply

Cyber Evaluations Engineer

Anthropic

San Francisco, CA | Washington, DCOn-siteFull-time

About this role

<div class="content-intro"><h2><strong>About Anthropic</strong></h2> <p>Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.</p></div><h2><strong>About the role</strong></h2> <p>We&39;re hiring Cyber Evaluations Engineers to build and run the evaluations that measure cyber-relevant capabilities and safeguard robustness in our models. You&39;ll design new evals, run per-release robustness testing, and dig into data on jailbreaks and prompt bypasses to understand where our safeguards hold up and where they don&39;t. You&39;ll also design many of the probes that detect cyber abuse in production and help shape the overall detection architecture alongside the policy team. </p> <h2><strong>Key responsibilities</strong></h2> <ul> <li> <p>Design and run capability, uplift, and safety evaluations to assess cyber-relevant risk in new models</p> </li> <li> <p>Execute per-release safeguard-robustness testing ahead of major model launches</p> </li> <li> <p>Analyze evaluation results and communicate findings clearly to the team and to stakeholders</p> </li> <li> <p>Design, prototype, and tune detection probes for cyber misuse</p> </li> <li> <p>Work with the cyber policy team to turn policy lines into a layered, robust abuse-detection architecture, and measure its precision and coverage over time</p> </li> <li> <p>Build and maintain internal tooling used to run and score evaluations</p> </li> <li> <p>Collaborate with policy and e

Related opportunities