ML Infrastructure Engineer
Nebius
RemoteRemote OKFull-time$140k – $220k / year
About this role
ML Infrastructure Engineer at Nebius - Remote
Nebius
Nebius enables B2B companies to build local hyperscaling cloud platforms with cost-effective GPUs, InfiniBand network, and 50% less compute cost. They offer managed Kubernetes and a launch-ready business model for innovative cloud solutions.
Internet Software & Services
Information Technology
51-250 (120)
158 open positions
Tags
E-commerce Computer Programming Software Professional Services Computers Information Technology & Services B2C B2B
Links
ML Infrastructure Engineer
2 months ago
Europe, United States
Senior
DevOps and Infrastructure
Nebius
Nebius enables B2B companies to build local hyperscaling cloud platforms with cost-effective GPUs, InfiniBand network, and 50% less compute cost. They offer managed Kubernetes and a launch-ready business model for innovative cloud solutions.
Internet Software & Services
51-250
LinkedIn View All Jobs 158
Description
- Work with hardware and development teams to profile and analyze GPU performance at the system and kernel level.
- Evaluate and compare GPU performance across different platforms, architectures, and software stacks such as CUDA and ROCm.
- Debug and optimize ML workloads to run efficiently on GPU hardware by identifying and resolving performance bottlenecks.
- Perform acceptance testing for new GPU clusters to verify performance, stability, and compatibility for AI workloads.
- Run experiments across diverse GPU system configurations to assess the impact of interconnect strategies and system-level optimizations.
- Develop tools and dashboards to visualize performance metrics, bottlenecks, and trends.
- Contribute to internal tooling, frameworks, and best practices for GPU benchmarking and optimization.
Requirements
- A strong understanding of the theoretical foundations of machine learning.
- Deep understanding of performance aspects of large neural network training and inference, including data, tensor, contex