Senior Cloud Reliability Engineer – AI-Powered SRE

Cerebras

Mountain View (CA)

On-site

USD 190,000 - 240,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cerebras Systems seeks a Staff Cloud Reliability Engineer to own the availability and efficiency of its multi-cloud SaaS platform (AWS, GCP). You will lead SRE practices with AI-powered automation, design scalable Kubernetes controller frameworks, and manage geodistributed data services such as PostgreSQL/RDS, MySQL and Elasticsearch/OpenSearch.

Expect on-call rotation and a strong bias toward automation. The role emphasizes capacity planning, incident response, and blameless post-mortems, with

Qualifications

  • B.Tech. degree in Computer Science or equivalent.
  • 6+ years of Enterprise SaaS Ops experience.
  • Proficiency in Go and Python; experience with Terraform and Ansible.
  • Expertise in Cloud Security and/or Cloud networking.
  • Experience with AI Ops tools and Agentic LLM.
  • Experience with cloud services, Kubernetes, cloud databases (PostgreSQL/RDS/MySQL/DynamoDB), Elasticsearch/OpenSearch, and caching tech.
  • Experience with observability (metrics, logs, tracing) and alerting in a Cloud SaaS environment.
  • Strong debugging and problem-solving skills across network, systems, database, and application domains.
  • Advanced cloud provider certifications (AWS, Azure, GCP) are a bonus; Full stack dev experience is a bonus.

Responsibilities

  • Operate a high-scale, multi-cloud SaaS platform (AWS, GCP) ensuring reliability, performance, and uptime for production workloads.
  • Embed AI-powered workflows into SRE practices: AI Ops platforms, autonomous agents for anomaly detection, triage, and runbook execution.
  • Lead capacity planning and scaling, modeling growth and implementing autoscaling to support SaaS growth without regressions.
  • Architect and operate Kubernetes controller frameworks; enforce cluster lifecycle, scheduling, autoscaling, and failover standards.
  • Own operations of cloud-native databases and data infrastructure (PostgreSQL/RDS, DynamoDB, MySQL, Elasticsearch/OpenSearch, ElastiCache) including performance tuning and backups.
  • Lead incident response and blameless post-mortems; drive root cause analysis for permanent fixes and prevention.
  • Promote automation-first culture with self-healing systems and GitOps (Terraform, Helm).
  • Participate in on-call rotations and act as senior escalation point during high-severity incidents.
  • Achieve measurable SaaS operational excellence via SLIs/SLOs/SLAs.

Skills

Go
Python
Terraform
Ansible
Kubernetes
Cloud security
AI Ops
LLM
Observability
Distributed systems

Education

B.Tech in Computer Science

Tools

PostgreSQL
RDS
MySQL
DynamoDB
Elasticsearch/OpenSearch
Redis
Kafka

Job description

Cerebras Systems seeks a Staff Cloud Reliability Engineer to own the availability and efficiency of its multi-cloud SaaS platform (AWS, GCP). You will lead SRE practices with AI-powered automation, design scalable Kubernetes controller frameworks, and manage geodistributed data services such as PostgreSQL/RDS, MySQL and Elasticsearch/OpenSearch.

Expect on-call rotation and a strong bias toward automation. The role emphasizes capacity planning, incident response, and blameless post-mortems, with

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI-Driven SRE for Cloud Reliability
Senior AI-Driven SRE for Cloud Reliability

Cerebras • Mountain View (CA)

Hybrid
USD 100,000 - 150,000
Competitive salary and benefits package
Opportunities for professional growth
Collaborative work environment
Senior Staff Engineer - AI Inference & Resilient Cloud
Senior Staff Engineer - AI Inference & Resilient Cloud

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 120,000 - 160,000
Inclusive work environment
Job stability with startup vitality
Opportunity for continuous learning
Senior/Staff Cloud Reliability Engineer
Senior/Staff Cloud Reliability Engineer

Cerebras • Mountain View (CA)

On-site
USD 190,000 - 240,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Lead Cloud SRE & AI-Driven Infra Architect
Lead Cloud SRE & AI-Driven Infra Architect

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Employee benefits
Diverse workplace
AI IT SRE Team Lead — Automation & Observability
AI IT SRE Team Lead — Automation & Observability

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 130,000 - 170,000
Diverse and inclusive work environment
Opportunity to work on cutting-edge AI research
Job stability with startup vitality
Senior SRE: AI-Driven Reliability & Customer Support
Senior SRE: AI-Driven Reliability & Customer Support

Cerebras • Mountain View (CA)

On-site
USD 90,000 - 120,000
Competitive salary and benefits package
Opportunities for professional growth
Collaborative work environment
Senior Staff Cloud Reliability Engineer for AI Infra
Senior Staff Cloud Reliability Engineer for AI Infra

Epoch Biodesign • San Francisco (CA)

On-site
USD 180,000 - 220,000
Health insurance
401(k) with employer match
Paid Parental Leave
+2
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Staff Site Reliability Engineer - Automation and Platform
Staff Site Reliability Engineer - Automation and Platform

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Job stability
Startup vitality
Non-corporate culture