Senior AI Compute Cluster Operations Engineer

Cerebras

Toronto

On-site

CAD 120,000 - 190,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerebras Systems seeks an AI Cluster Operations Engineer to manage the world's largest AI compute clusters, including the Wafer-Scale Engine. You will ensure health, performance, and availability of infrastructure, maximize capacity, and support AI initiatives.

The role requires Linux expertise, Docker/Kubernetes, and experience with monitoring, automation, and distributed systems. You will work with a fast-paced, on-call team to deliver reliable, scalable platforms for cutting-edge AI workloads.

Qualifications

  • 6-8 years of experience managing and operating complex compute infrastructure, preferably with ML/HPC context.
  • Proficient in Python and Go, with experience building operational platforms and reliability tooling.
  • Deep knowledge of Linux-based compute systems and CLI tools.
  • Extensive Docker experience and orchestration via Kubernetes.
  • Experience with monitoring and alerting systems and troubleshooting complex issues.

Responsibilities

  • Deploy, configure, and debug container-based services using Docker.
  • Build and own software that powers cluster operations, including dashboards and reliability tooling.
  • Collaborate with cross-functional teams to translate requirements into scalable OM platforms.
  • Develop APIs, automation services, and integrations for visibility and fleet management.
  • Manage and operate multiple advanced AI compute clusters.
  • Monitor cluster health and proactively resolve issues to maximize compute capacity.
  • Provide 24/7 monitoring and hands-on troubleshooting as needed.
  • Handle escalations and collaborate to resolve complex challenges.
  • Stay up-to-date with advancements in AI compute infrastructure.

Skills

6-8 years experience
Python
Go
Distributed systems
Linux
Docker
Kubernetes
Troubleshooting
Monitoring/alerts
On-call rotation
Communication

Tools

Docker
Kubernetes

Job description

Cerebras Systems seeks an AI Cluster Operations Engineer to manage the world's largest AI compute clusters, including the Wafer-Scale Engine. You will ensure health, performance, and availability of infrastructure, maximize capacity, and support AI initiatives.

The role requires Linux expertise, Docker/Kubernetes, and experience with monitoring, automation, and distributed systems. You will work with a fast-paced, on-call team to deliver reliable, scalable platforms for cutting-edge AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Cluster Operations Engineer
Senior AI Cluster Operations Engineer

Cerebras Systems • Toronto

On-site
CAD 120,000 - 160,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras Systems • Toronto

On-site
CAD 120,000 - 160,000
Distributed Software Engineer
Distributed Software Engineer

Cerebras Systems, Inc. • Ottawa

On-site
CAD 90,000 - 120,000
Job stability with startup vitality
Open access to cutting-edge AI research
Senior Software Development Engineer in Test (SDET) - AI Cluster
Senior Software Development Engineer in Test (SDET) - AI Cluster

Cerebras Systems, Inc. • Toronto

On-site
CAD 140,000 - 210,000
Senior Software Development Engineer in Test (SDET) - AI Cluster
Senior Software Development Engineer in Test (SDET) - AI Cluster

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Toronto

On-site
CAD 170,000 - 210,000
ML Systems Integration Engineer
ML Systems Integration Engineer

Cerebras • Toronto

On-site
CAD 90,000 - 150,000
HPC Engineer
HPC Engineer

Wyatt Partners • Toronto

On-site
CAD 90,000 - 130,000
DevOps Engineer - New Grad 2026
DevOps Engineer - New Grad 2026

Cerebras Systems, Inc. • Toronto

On-site
CAD 70,000 - 90,000
Opportunity to work on an innovative AI platform
Diverse and inclusive work environment
Job stability with startup vitality