Principal SRE - AI Inference

Cerebras

Sunnyvale (CA)

On-site

USD 260,000 - 380,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cerebras Systems is seeking a Principal SRE to define and drive the architecture for scaling our AI inference fleet across datacenters and cloud-based solutions. You will build self-service platforms, observability, and automation, enabling product teams and customers to operate with strong guardrails.

In the first year, you will lead a transformation from ops-centric reliability to a shared engineering discipline, mentor senior engineers, and shape capacity management and rollout safety.

Qualifications

  • 15+ years in SRE or related infra, with large-scale reliability improvements.
  • Experience building production control planes, schedulers, and orchestration.
  • Ability to lead complex programs and mentor senior engineers.
  • Proven cross-functional collaboration with product and customers.
  • Hands-on with observability, incident response, and postmortems.

Responsibilities

  • Define and deliver scalable, reliable infra across multi-datacenter and cloud.
  • Architect self-service platforms with safe workflows.
  • Establish SLOs/SLIs, error budgets, and capacity forecasting.
  • Mentor senior SREs and drive automation of toil.
  • Track impact with deployment velocity and MTTR metrics.

Skills

SRE
Infrastructure engineering
Platform engineering
Observability
Incident response
SLOs/SLIs
Cross-team leadership

Tools

Bazel

Job description

About the Role

We are building a high-performance SRE function to support one of the world’s fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE). This team will help deliver world-class, ultra-reliable inference infrastructure for leading model builders such as OpenAI and other frontier labs.

As a Principal SRE, you will define and drive the technical architecture for scaling our inference fleet through self-service delivery, shared observability, capacity orchestration, rollout safety, and operational automation. This role starts with 2–3 weeks of hands-on operational immersion to build deep context on the current stack, production pain points, and high-stakes workflows.

From there, your mandate shifts to architecting the “tomorrow” layer: a unified capacity management and production control plane that enables reliable capacity planning, workload placement, rollout safety, validation, and operational decision-making across large-scale inference infrastructure.

Success in the first year means core engineering teams, product managers, external customers, and cluster stakeholders can execute critical operational workflows through self-service systems with strong guardrails, clear ownership, and minimal dependency on expert SRE operators.

You will collaborate with the tech leads and the leadership team across core, cluster, cloud, and product stakeholders. This work will shift reliability from an ops-only burden to a shared engineering discipline that underpins frontier AI inference at scale.

If you are a proven Principal engineer who enjoys turning complexity into elegant reliability at scale, this is your chance to lead this transformation from the front.

This role does not require 24/7 on-call rotations.

Key Responsibilities
  • Define and implement a robust strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions.
  • Architect self-service platforms and internal tooling that let product teams, external customers, and cluster operators safely trigger and observe critical workflows with minimal handoffs.
  • Define and evolve reliability practices for inference workloads, including SLOs and SLIs for latency, throughput, and accuracy stability; error budgets; blameless postmortems; chaos testing; and capacity forecasting across multi-datacenter and on-prem environments.
  • Mentor senior SREs, support critical incident escalations, and use production pain points to prioritize the highest-leverage automation work.
  • Measure and drive impact through clear metrics, including toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
Required Experience & Skills
  • 15+ years in SRE, infrastructure engineering, or platform engineering, with a record of setting technical direction and delivering reliability improvements at large scale in FAANG, hyperscaler, frontier AI, or similarly demanding production environments.
  • Deep experience with large-scale compute fleets, internal control planes, schedulers, orchestration systems, capacity management, and reliability automation.
  • Experience defining and driving cross-team architecture for production control planes, capacity orchestration, fleet management, or self-service infrastructure platforms with clear operational ownership.
  • Strong judgment in converging fragmented workflows, tools, and teams into coherent architectures that improve reliability, efficiency, and operational leverage.
  • Ability to lead complex, ambiguous technical programs end to end; influence senior cross-functional stakeholders; mentor senior engineers; and communicate technical strategy clearly.
  • Hands-on experience with production observability, incident response, and SLO-based reliability management across metrics, logs, traces, alerting, dashboards, and operational review loops.
Nice-to-Haves
  • Experience with Bazel or other large-scale build systems in production.
  • Background in AI/ML inference systems, including model serving runtimes, disaggregated inference, GPU orchestration, latency and accuracy SLOs, or drift monitoring.
  • Prior work on predictive autoscaling, chaos engineering, or cost-aware capacity management for compute-intensive workloads.
Location
  • SF Bay Area
  • Toronto
Why Join Cerebras
  1. Build a breakthrough AI platform beyond the constraints of the GPU.
  2. Publish and open source their cutting-edge AI research.
  3. Work on one of the fastest AI supercomputers in the world.
  4. Enjoy job stability with startup vitality.
  5. Our simple, non-corporate work culture that respects individual beliefs.

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click to review our CCPA disclosure notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Ops & Automation
Site Reliability Engineer - Ops & Automation

Cerebras • United States

On-site
USD 125,000 - 170,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
IT SRE Team Lead
IT SRE Team Lead

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 130,000 - 170,000
Diverse and inclusive work environment
Opportunity to work on cutting-edge AI research
Job stability with startup vitality
Principal Engineer, Inference Cloud
Principal Engineer, Inference Cloud

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Opportunity to work on cutting-edge AI technology
Inclusive and supportive work environment
Job stability with startup vitality
Principal Engineer, AI Inference Reliability
Principal Engineer, AI Inference Reliability

Cerebras • United States

On-site
USD 120,000 - 160,000
Inclusive work environment
Opportunities for continuous learning
Startup vitality with job stability
Senior SDET, Inference Platform
Senior SDET, Inference Platform

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 190,000
Principal Engineer, Inference Cloud
Principal Engineer, Inference Cloud

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Job stability with startup vitality
Open-source AI research
Simple, non-corporate work culture
AI Inference Core - Software Integration Engineer
AI Inference Core - Software Integration Engineer

Cerebras • United States

Hybrid
USD 150,000 - 190,000
AI Inference Core - SDET Technical Lead, Release Integration Testing
AI Inference Core - SDET Technical Lead, Release Integration Testing

Cerebras Systems • Sunnyvale (CA)

Hybrid
CAD 120,000 - 180,000
IT SRE Team Lead
IT SRE Team Lead

Cerebras • Sunnyvale (CA)

On-site
USD 130,000 - 160,000