Senior AI Cluster Operations Engineer

Cerebras Systems

Toronto

On-site

CAD 120,000 - 160,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerebras Systems is seeking an AI Cluster Operations Engineer to manage and operate our Wafer-Scale Engine compute clusters, ensuring health, performance, and availability to support cutting-edge AI workloads.

You will develop automation, monitoring dashboards, and reliability tooling, work with Linux-based systems, Docker and Kubernetes, and provide 24/7 on-call support across global infrastructure to maximize compute capacity.

Qualifications

  • 6-8 years of experience managing and operating complex compute infrastructure.
  • Proficiency in Python and Go for building operational platforms.
  • Experience with distributed systems and Linux-based compute environments.
  • Extensive knowledge of Docker and container orchestration (Kubernetes).

Responsibilities

  • Deploy, configure, and debug container-based services using Docker.
  • Build and own software for cluster operations including monitoring platforms and reliability tooling.
  • Collaborate with cross-functional teams to translate requirements into scalable O&M products.
  • Develop APIs, automation services, and integrations for global AI infrastructure.
  • Monitor cluster health and maximize compute capacity.
  • Provide 24/7 monitoring and hands-on troubleshooting.

Skills

Python
Go
Distributed systems
Linux
Docker
Kubernetes

Job description

Cerebras Systems is seeking an AI Cluster Operations Engineer to manage and operate our Wafer-Scale Engine compute clusters, ensuring health, performance, and availability to support cutting-edge AI workloads.

You will develop automation, monitoring dashboards, and reliability tooling, work with Linux-based systems, Docker and Kubernetes, and provide 24/7 on-call support across global infrastructure to maximize compute capacity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Compute Cluster Operations Engineer
Senior AI Compute Cluster Operations Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras Systems • Toronto

On-site
CAD 120,000 - 160,000
Distributed Software Engineer
Distributed Software Engineer

Cerebras Systems, Inc. • Ottawa

On-site
CAD 90,000 - 120,000
Job stability with startup vitality
Open access to cutting-edge AI research
Senior Software Development Engineer in Test (SDET) - AI Cluster
Senior Software Development Engineer in Test (SDET) - AI Cluster

Cerebras Systems, Inc. • Toronto

On-site
CAD 140,000 - 210,000
Senior Software Development Engineer in Test (SDET) - AI Cluster
Senior Software Development Engineer in Test (SDET) - AI Cluster

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Toronto

On-site
CAD 170,000 - 210,000
ML Systems Integration Engineer
ML Systems Integration Engineer

Cerebras • Toronto

On-site
CAD 90,000 - 150,000
Next-Gen AI Performance Engineer - Kernel & WSE Optimization
Next-Gen AI Performance Engineer - Kernel & WSE Optimization

Cerebras Systems • Toronto

On-site
CAD 120,000 - 180,000
DevOps Engineer - New Grad 2026
DevOps Engineer - New Grad 2026

Cerebras Systems, Inc. • Toronto

On-site
CAD 70,000 - 90,000
Opportunity to work on an innovative AI platform
Diverse and inclusive work environment
Job stability with startup vitality