Apply

AI Infrastructure Systems Engineer

Together AI

San FranciscoOn-siteFull-time

About this role

<h3><strong>About the Role</strong></h3> <p>At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale.</p> <p>If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk.</p> <h3><strong>Responsibilities</strong></h3> <ul> <li>Design and build <strong>fleet automation systems</strong> that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.</li> <li>Build <strong>AI Infrastructure Agents</strong> that automate deployment, root-cause failures, incident triage, and autonomous remediation.</li> <li>Develop <strong>Fleet Intelligence</strong> platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers.</li> <li>Build software that maximizes <strong>GPU availability, utilization, performance, and reliability</strong> across thousands of accelerators.</li> <li>Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.</li> <li>Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations.</li> <li>Continuously improve deployment velocity, reliability, and operational efficiency through automation.</li> <li>Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure.&

Related opportunities