AI Infrastructure Engineer, Serving Platform
Scale AI
London, UKOn-siteFull-time
About this role
<p>As a Software Engineer on the ML Infrastructure team, you will design and build platforms for scalable, reliable, and efficient serving of LLMs. Our platform powers cutting-edge research and production systems, supporting both internal and external use cases across various environments.</p> <p>&nbsp;</p> <p>The ideal candidate combines strong ML fundamentals with deep expertise in backend system design. You’ll work in a highly collaborative environment, bridging research and engineering to deliver seamless experiences to our customers and accelerate innovation across the company.</p> <h2>You will:</h2> <ul> <li>Build and maintain fault-tolerant, high-performance systems for serving LLMs and other models at scale.</li> <li>Build an internal platform to empower LLM capability discovery.</li> <li>Collaborate with researchers and engineers to integrate and optimize models for production and research use cases.</li> <li>Conduct architecture and design reviews to uphold best practices in system design and scalability.</li> <li>Develop monitoring and observability solutions to ensure system health and performance.</li> <li>Lead projects end-to-end, from requirements gathering to implementation, in a cross-functional environment.</li> </ul> <h2>Ideally you&39;d have:</h2> <ul> <li>4+ years of experience building large-scale, high-performance backend systems.</li> <li>Strong programming skills in one or more languages (e.g., Python, Go, Rust, C++).</li> <li>Experience with LLM serving and routing fundamentals (e.g. rate limiting, token streaming, load balancing, budgets, etc.)</li> <li>Experience with LLM capabilities and concepts such as reasoning, tool calling, prompt templates, etc.</li> <li>Experience with containers and orchestration tools (e.g., Docker, Kuberne