AI Compute Cluster Engineer — Scale & Reliability

Cerebras

San Francisco (CA)

On-site

CAD 209,000 - 321,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Cerebras Systems is seeking an AI Cluster Operations Engineer to manage and optimize our Wafer-Scale Engine compute clusters and related infrastructure. You will ensure health, performance, and availability while expanding capacity for fast-moving AI initiatives.

You will develop APIs and automation, deploy Docker-based services, monitor fleets, and collaborate across teams to deliver scalable O&M products. A strong Linux background and 6–8 years in infra are required.

Qualifications

  • 6-8 years of relevant experience in managing and operating complex compute infrastructure.
  • Proficient in Python and Go, with experience building operational platforms and tooling.
  • Experience and expertise in distributed systems is a must.
  • Deep understanding of Linux-based compute systems and CLI tools.
  • Extensive knowledge of Docker containers and container orchestration platforms like Kubernetes.
  • Proven ability to troubleshoot and resolve complex technical issues.
  • Experience with monitoring and alerting systems.
  • Willingness to participate in a 24/7 on-call rotation.

Responsibilities

  • Deploy, configure, and debug container-based services using Docker.
  • Build and own software solutions that power cluster operations and reliability tooling.
  • Collaborate with cross-functional teams to translate operational requirements into scalable O&M products.
  • Develop APIs, automation services, and integrations for fleet management and incident response.
  • Manage and operate multiple advanced AI compute infrastructure clusters.
  • Monitor and oversee cluster health, proactively identifying issues.
  • Maximize compute capacity through optimization and resource allocation.
  • Provide 24/7 monitoring and hands-on troubleshooting.
  • Handle engineering escalations and collaborate with other teams to resolve complex challenges.
  • Stay up-to-date with advancements in AI compute infrastructure.

Skills

Python
Go
Distributed systems
Linux
Monitoring & alerting
On-call readiness
Collaboration
Communication

Tools

Docker
Kubernetes

Job description

Cerebras Systems is seeking an AI Cluster Operations Engineer to manage and optimize our Wafer-Scale Engine compute clusters and related infrastructure. You will ensure health, performance, and availability while expanding capacity for fast-moving AI initiatives.

You will develop APIs and automation, deploy Docker-based services, monitor fleets, and collaborate across teams to deliver scalable O&M products. A strong Linux background and 6–8 years in infra are required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Compute Platform Architect — AI Cluster Scale
Compute Platform Architect — AI Cluster Scale

Cerebras • United States

On-site
USD 150,000 - 200,000
Equity options
Open-source contributions
Inclusive work culture
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras • San Francisco (CA)

On-site
CAD 209,000 - 321,000
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Data Center Infra Engineer - High-Density Compute
AI Data Center Infra Engineer - High-Density Compute

Cerebras Systems • Sunnyvale (CA)

On-site
USD 120,000 - 190,000
Flexible hardware design projects
AI Data Center Infra Engineer: Deploy & Automate Scale
AI Data Center Infra Engineer: Deploy & Automate Scale

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
AI Data Center Infrastructure Architect
AI Data Center Infrastructure Architect

Cerebras Systems • California (MO)

On-site
USD 120,000 - 190,000
AI Compute Reliability Engineer — Scale & Incident Leadership
AI Compute Reliability Engineer — Scale & Incident Leadership

Fluidstack • Seattle (WA)

On-site
USD 173,000 - 224,000
Global AI Data Center Deployment Lead
Global AI Data Center Deployment Lead

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 200,000
Safety program
AI Security Leader for Large-Scale Clusters
AI Security Leader for Large-Scale Clusters

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 240,000
Software Engineer, Cluster Deployment
Software Engineer, Cluster Deployment

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 190,000