Senior Distributed Systems Engineer — Kubernetes & AI Infra

Engg

Toronto

On-site

CAD 140,000 - 210,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cerebras Systems is seeking a seasoned cluster engineer to join the team responsible for turning thousands of wafers, servers, and switches into a reliable, observable cloud. The role focuses on automating bare-metal provisioning, upgrades, and security patches, with a strong emphasis on CRDs, controllers, and RBAC in a Go/Python/Kubernetes stack.

You will design and operate large inference workloads, build telemetry pipelines with Prometheus and Grafana, and contribute to resilient,

Qualifications

  • 5+ years building and operating production distributed systems or infrastructure software.
  • Production-quality Go and Python.
  • Real Kubernetes depth: controllers, CRDs, RBAC, and reconciliation semantics.
  • Strong debugging skills across distributed systems, Linux, and networking.
  • Prometheus and Grafana experience; metrics and exporter design.

Responsibilities

  • Declarative, CRD-driven automation of bare-metal networking, OS, and application software across clusters of Cerebras systems, servers, and switches, built to reconcile thousands of nodes.
  • Push-button cluster install, upgrade, and security patching with real downtime budgets, gated by canaries.
  • Kubernetes operators that schedule large inference workloads: resource locks, priority queues, topology placement, health-aware placement.
  • gRPC control-plane services, authorization, admission webhooks, and quota policy for a multi-tenant fleet.
  • Metrics and log pipelines with exporters for wafer-scale systems, servers, and network fabric on Prometheus and Grafana, with SLOs and alerting.
  • Failure detection, HA control planes, and automated recovery, plus the CLIs, APIs, and MCP gateway that expose the fleet to users, operators, and AI agents.

Skills

distributed systems
Go & Python
Kubernetes depth
debugging distributed systems
Prometheus & Grafana
self-driving capability
AI tooling adoption
infrastructure automation

Job description

Cerebras Systems is seeking a seasoned cluster engineer to join the team responsible for turning thousands of wafers, servers, and switches into a reliable, observable cloud. The role focuses on automating bare-metal provisioning, upgrades, and security patches, with a strong emphasis on CRDs, controllers, and RBAC in a Go/Python/Kubernetes stack.

You will design and operate large inference workloads, build telemetry pipelines with Prometheus and Grafana, and contribute to resilient,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Compute Cluster Operations Engineer
Senior AI Compute Cluster Operations Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Distributed Software Engineer
Distributed Software Engineer

Cerebras Systems • Toronto

On-site
CAD 140,000 - 210,000
Distributed Software Engineer
Distributed Software Engineer

Engg • Toronto

On-site
CAD 140,000 - 210,000
Sr. Staff Software Engineer, Inference Platform
Sr. Staff Software Engineer, Inference Platform

Cerebras Systems • Lower Sackville

On-site
CAD 180,000 - 240,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras Systems • Lower Sackville

On-site
CAD 250,000 - 333,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Toronto

On-site
CAD 170,000 - 210,000
Senior Software Development Engineer in Test (SDET) - AI Cluster
Senior Software Development Engineer in Test (SDET) - AI Cluster

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Senior Software Development Engineer in Test (SDET) - AI Cluster
Senior Software Development Engineer in Test (SDET) - AI Cluster

Cerebras Systems, Inc. • Toronto

On-site
CAD 140,000 - 210,000
Senior SDET, Inference Platform
Senior SDET, Inference Platform

Cerebras • Toronto

On-site
CAD 120,000 - 160,000