Senior Reliability Engineer, AI Inference at Scale

Cerebras

Raleigh (NC)

On-site

USD 170,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerebras Systems is seeking a hands-on Reliability Tech Lead to own the mission of making Cerebras Inference the most reliable AI service. You will establish SLOs and drive reliability across the inference stack, from client SDKs to data-center deployments.

You will define incident-response frameworks, design scalable reliability mechanisms, and collaborate with hundreds of engineers to uphold world-class reliability standards.

Qualifications

  • 7+ years of backend, infrastructure, or reliability engineering for large-scale distributed systems.
  • Strong programming skills in Python, C++, Go, or Rust.
  • Deep experience with reliability: SLO/SLI/SLA design, incident response, postmortem culture.
  • Excellent communication and cross-functional leadership skills.
  • Bonus: experience building large-scale AI infrastructure systems.

Responsibilities

  • Define and drive reliability strategy: establish SLOs and ensure alignment across engineering.
  • Design and implement reliability mechanisms: build, evolve systems for fault detection, degradation, failover, throttling, and recovery across multiple regions and data centers.
  • Lead large-scale incident management: own postmortems, root-cause analysis, and prevention loops for reliability incidents.
  • Architect for reliability and observability: influence system design for redundancy, durability, and debuggability.
  • Develop reliability tooling: create internal tools and frameworks for chaos testing, load simulation, and distributed fault injection.
  • Collaborate broadly: work across software, infrastructure, and hardware teams to embed reliability across layers of the inference service.
  • Monitor and communicate reliability metrics: build dashboards and alerts for service health and actionable insights.
  • Mentor and influence: guide engineers and share best practices for designing, testing, and operating reliable large-scale systems.

Skills

Backend engineering
Distributed systems
SLO/SLI design
Incident response
Leadership
Python
C++
Go
Rust

Education

Bachelor's degree in CS
Master's degree in CS

Job description

Cerebras Systems is seeking a hands-on Reliability Tech Lead to own the mission of making Cerebras Inference the most reliable AI service. You will establish SLOs and drive reliability across the inference stack, from client SDKs to data-center deployments.

You will define incident-response frameworks, design scalable reliability mechanisms, and collaborate with hundreds of engineers to uphold world-class reliability standards.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Reliability Engineer
Senior AI Inference Reliability Engineer

Cerebras • United States

On-site
USD 120,000 - 160,000
Inclusive work environment
Opportunities for continuous learning
Startup vitality with job stability
Lead AI Inference Reliability Engineer
Lead AI Inference Reliability Engineer

Cerebras Systems • United States

Hybrid
USD 120,000 - 180,000
Job stability with startup vitality
Simple, non-corporate work culture
Opportunity to work on AI supercomputers
Principal Engineer, AI Inference Reliability
Principal Engineer, AI Inference Reliability

Cerebras • Raleigh (NC)

On-site
USD 170,000 - 250,000
Principal Engineer, AI Inference Reliability
Principal Engineer, AI Inference Reliability

Cerebras Systems • United States

Hybrid
USD 120,000 - 180,000
Job stability with startup vitality
Simple, non-corporate work culture
Opportunity to work on AI supercomputers
Principal Engineer, AI Inference Reliability
Principal Engineer, AI Inference Reliability

Cerebras • United States

On-site
USD 120,000 - 160,000
Inclusive work environment
Opportunities for continuous learning
Startup vitality with job stability
Principal SRE: AI Inference Reliability & Scale
Principal SRE: AI Inference Reliability & Scale

Cerebras • Sunnyvale (CA)

On-site
USD 260,000 - 380,000
Senior Staff Engineer - AI Inference & Resilient Cloud
Senior Staff Engineer - AI Inference & Resilient Cloud

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 120,000 - 160,000
Inclusive work environment
Job stability with startup vitality
Opportunity for continuous learning
Principal SRE, AI Inference — Scale & Self-Service
Principal SRE, AI Inference — Scale & Self-Service

Cerebras Systems • Sunnyvale (CA)

Hybrid
USD 200,000 - 320,000
Staff Software Engineer, AI Inference Platform
Staff Software Engineer, AI Inference Platform

Cerebras • Sunnyvale (CA)

On-site
USD 140,000 - 200,000
Senior SDET - AI Inference Platform Quality & Automation
Senior SDET - AI Inference Platform Quality & Automation

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 190,000