Distributed Software Engineer
Cerebras
Toronto, CANRemote OKFullTime
About this role
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. The Role The Cluster engineering team owns the software that turns thousands of wafers, servers, and switches into a cloud that stays up, stays busy, and stays debuggable. We stand clusters up from bare metal, schedule training and inference workloads across the fleet, keep it healthy, and make it observable to users, operators, and increasingly to AI agents. The stack is Go and Python on Kubernetes, running both on-premise deployments and our own cloud. Responsibilities · Declarative, CRD-driven automation of bare-metal networking, OS, and application software across clusters of Cerebras systems, servers, and switches, built to reconcile thousands of nodes · Push-button cluster install, upgrade, and security patching with real downtime budgets, gated by canaries · Kubernetes operators that schedule large inference workload: resource locks, priority queues, network topology, and health-aware placement · gRPC control-plane services, authorization, admission webhooks, and quota policy for a multi-tenant fleet · Metrics and log pipelines with purpose-built exporters for wafer-scale systems, servers (Redfish, IPMI), and network fabric (gNMI, sFlow), on Prometheus and Grafana, with SLOs and alerting · Failure detection, HA control planes, and automated recovery, plus the CLIs, APIs, an