Apply

Senior Software Engineer - Together Cloud Infrastructure

Together AI

San FranciscoOn-siteFull-time

About this role

<h3>About the Role</h3> <p>Together AI is building the AI Native Cloud, an end-to-end platform for the full<br>generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art<br>AI cloud infrastructure. The Together Cloud team builds the Together GPU Clusters<br>product, which provides high-performance, AI-ready GPU clusters through a self-serve<br>cloud console and is the virtualized infrastructure layer powering Together’s inference,<br>RL, and fine-tuning products.</p> <p><br>As a Senior Software Engineer in Together Cloud Infrastructure, you will play a key role<br>in building the next generation AI cloud platform – a highly available, global, blazing-fast<br>cloud infrastructure that virtualizes cutting-edge ML hardware (GB200s/GB300s,<br>BlueField DPUs). You&39;ll enable state-of-the-art ML practitioners with self-serve AI cloud<br>services, such as on-demand + managed Kubernetes and Slurm clusters, for both our<br>internal SaaS products (inference, fine-tuning, RL) and our external cloud customers,<br>spanning dozens of data centers across the world.</p> <h3><strong>Responsibilities</strong></h3> <ul> <li>Design, build, and maintain performant, secure, and highly-available backend services/operators that run in our data centers and automate hardware management, such as Infiniband partitioning, in. DC parallel storage provisioning, and VM provisioning.</li> <li>Design and build out the IaaS software layer for a new GB200 data center with thousands of GPUs.</li> <li>Design and build distributed GPU scheduling and the global management plane that power on-demand and managed clusters across dozens of data centers</li> <li>Develop infrastructure that powers our internal inference, RL, and fine-tuning product

Related opportunities