Senior/Staff Cloud Reliability Engineer

Cerebras

Mountain View (CA)

On-site

USD 190,000 - 240,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerebras Systems seeks a Staff Cloud Reliability Engineer to own the availability and efficiency of its multi-cloud SaaS platform (AWS, GCP). You will lead SRE practices with AI-powered automation, design scalable Kubernetes controller frameworks, and manage geodistributed data services such as PostgreSQL/RDS, MySQL and Elasticsearch/OpenSearch.

Expect on-call rotation and a strong bias toward automation. The role emphasizes capacity planning, incident response, and blameless post-mortems, with

Qualifications

  • B.Tech. degree in Computer Science or equivalent.
  • 6+ years of Enterprise SaaS Ops experience.
  • Proficiency in Go and Python; experience with Terraform and Ansible.
  • Expertise in Cloud Security and/or Cloud networking.
  • Experience with AI Ops tools and Agentic LLM.
  • Experience with cloud services, Kubernetes, cloud databases (PostgreSQL/RDS/MySQL/DynamoDB), Elasticsearch/OpenSearch, and caching tech.
  • Experience with observability (metrics, logs, tracing) and alerting in a Cloud SaaS environment.
  • Strong debugging and problem-solving skills across network, systems, database, and application domains.
  • Advanced cloud provider certifications (AWS, Azure, GCP) are a bonus; Full stack dev experience is a bonus.

Responsibilities

  • Operate a high-scale, multi-cloud SaaS platform (AWS, GCP) ensuring reliability, performance, and uptime for production workloads.
  • Embed AI-powered workflows into SRE practices: AI Ops platforms, autonomous agents for anomaly detection, triage, and runbook execution.
  • Lead capacity planning and scaling, modeling growth and implementing autoscaling to support SaaS growth without regressions.
  • Architect and operate Kubernetes controller frameworks; enforce cluster lifecycle, scheduling, autoscaling, and failover standards.
  • Own operations of cloud-native databases and data infrastructure (PostgreSQL/RDS, DynamoDB, MySQL, Elasticsearch/OpenSearch, ElastiCache) including performance tuning and backups.
  • Lead incident response and blameless post-mortems; drive root cause analysis for permanent fixes and prevention.
  • Promote automation-first culture with self-healing systems and GitOps (Terraform, Helm).
  • Participate in on-call rotations and act as senior escalation point during high-severity incidents.
  • Achieve measurable SaaS operational excellence via SLIs/SLOs/SLAs.

Skills

Go
Python
Terraform
Ansible
Kubernetes
Cloud security
AI Ops
LLM
Observability
Distributed systems

Education

B.Tech in Computer Science

Tools

PostgreSQL
RDS
MySQL
DynamoDB
Elasticsearch/OpenSearch
Redis
Kafka

Job description

Staff Cloud Reliability Engineer

We are seeking a Staff Site Reliability Engineer with deep enterprise SaaS operations expertise to own the availability, reliability, security, and efficiency of our Multi-Cloud (AWS, GCP) production SaaS platform. The ideal candidate brings hands‑on experience running highly available, large‑scale Kubernetes‑based control and data planes, a strong bias toward automation and AI‑augmented operations, and a proven track record in production security, capacity management, and cloud‑native data infrastructure.

Responsibilities
  • Operate a high‑scale, multi‑cloud (AWS, GCP) SaaS platform — ensuring reliability, performance, and uptime for business‑critical production workloads.
  • Embed AI and Agentic workflows into SRE practice: leverage AI Ops platforms and LLM‑powered autonomous agents for anomaly detection, automated triage, runbook execution, and incident summarization to reduce MTTR.
  • Drive capacity planning and scaling operations — proactively model growth, right‑size infrastructure, and implement horizontal/vertical autoscaling strategies to support SaaS growth without reliability regression.
  • Architect and operate Kubernetes controller frameworks governing both control plane and data plane services; define and enforce operational standards for cluster lifecycle, workload scheduling, autoscaling, and failover.
  • Own operations of high‑scale cloud‑native databases and data infrastructure: PostgreSQL/RDS, DynamoDB, MySQL, Elasticsearch/OpenSearch, ElastiCache (Redis/Memcached) on AWS and GCP — including performance tuning, backup/recovery, and incident response.
  • Lead incident response and blameless post‑mortems for P0/P1 events; drive root cause analysis to permanent resolution and prevention — eliminating repeat incidents through systemic fixes, not workarounds.
  • Define and enforce a culture of automation‑first: identify and eliminate toil through self‑healing systems, automated remediation pipelines, and infrastructure‑as‑code (Terraform, Helm, GitOps).
  • Participate in on‑call rotations for critical cloud infrastructure; serve as a senior escalation point and incident commander during high‑severity events.
  • Achieve quantifiable SaaS operational Excellence measured by related SLI/SLO/SLA
Required skills/qualifications
  • B.Tech. degree in Computer Science or equivalent.
  • At least 6+ years of Enterprise SaaS Ops experience
  • Strong proficiency in programming, particularly with Go and Python, and experience with Infrastructure as Code (IaC) tools like Terraform and Ansible.
  • Expertise in Cloud Security and/or Cloud networking
  • Experience with AI Ops tools, Agentic LLM.
  • Experience/ Knowledge in Cloud Services, Kubernetes, Cloud Databases like Postgres/RDS/MySQL/DynamoDB, Elastic, Kafka, and Microservice architecture is a bonus.
  • Experience in implementing and operating enterprise‑grade observability ( metrics, logs, tracing), alerting stack in a Cloud SaaS environment
  • Strong debugging and problem‑solving skills (network, systems, database, and application).
  • Advanced professional certifications from Cloud Providers ( AWS, Azure, GCP) in domains like K8s, Solution architecture, networking, and databases are a bonus.
  • Full Stack Architecture/Development Experience is a bonus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Cloud Operations Reliability Engineer (SRE)
Sr. Cloud Operations Reliability Engineer (SRE)

NextGen Healthcare • Georgia

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000
Senior Software Engineer – Cloud Platform & Operations
Senior Software Engineer – Cloud Platform & Operations

TalentCloud Group • Indiana

Hybrid
USD 100,000 - 130,000
Competitive salary
Equity
Flexible working hours
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veriipro • Charlotte (NC)

On-site
USD 140,000 - 190,000
Lead Cloud Platform Engineer / DevOps & SRE
Lead Cloud Platform Engineer / DevOps & SRE

Compunnel, Inc. • San Leandro (CA)

On-site
USD 150,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000
Senior Site Reliability Engineer NEX
Senior Site Reliability Engineer NEX

Patterson-UTI • Houston (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000