Junior/Senior or Staff Software Engineer, Inference / Compute Infrastructure Engineering
Together AI
India On-siteFull-time
About this role
<h3><strong>About the Role</strong></h3> <p><strong>REMOTE IN INDIA</strong></p> <div id="message-list_1787877133.837419"> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div>We&39;re looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You&39;ll design a manifest-driven API where the inference team declares what they need, whether that&39;s a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You&39;ll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we&39;re after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You&39;ll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered.</div> </div> </div> <