Senior SRE: AI Infra, On-Call, Kubernetes & Ceph

Engg

San Francisco (CA)

On-site

USD 180,000 - 230,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Engg, a San Francisco–based AI infrastructure company, seeks a senior on-call SRE to own incident response for Kubernetes, Ceph storage, and bare metal servers. You will coordinate with enterprise clients during outages, drive resolution timelines, and contribute to runbooks and automation.

You will lead cross-team investigations, manage etcd backups/restores, and help reduce incident frequency through reliability improvements and IaC tooling.

Qualifications

  • :years of production Kubernetes experience in enterprise environments.
  • Experience with distributed storage Ceph or equivalents.
  • Experience with at least one CNI plugin (Cilium/Calico).
  • Strong Linux admin and bare metal background.
  • Experience managing etcd backups/restores.

Responsibilities

  • Respond to and resolve production incidents across client infrastructure (Kubernetes, Ceph, bare metal).
  • Troubleshoot complex distributed systems issues during live outages.
  • Handle escalations requiring deep etcd, Ceph RGW, Cilium, and load balancer expertise.
  • Communicate incident status and timelines to enterprise clients.
  • Participate in follow-the-sun on-call rotation across time zones.
  • Document incidents and improve runbooks based on patterns.
  • Collaborate with infra on reliability and automation improvements.

Skills

Kubernetes operations
Distributed systems troubleshooting
On-call incident response
Linux administration
Client communication

Tools

Ceph
etcd
Cilium
NVIDIA Kubernetes Operator
Kubespray
Ansible

Job description

Engg, a San Francisco–based AI infrastructure company, seeks a senior on-call SRE to own incident response for Kubernetes, Ceph storage, and bare metal servers. You will coordinate with enterprise clients during outages, drive resolution timelines, and contribute to runbooks and automation.

You will lead cross-team investigations, manage etcd backups/restores, and help reduce incident frequency through reliability improvements and IaC tooling.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Infrastructure/Site Reliability Engineer On-Call
Senior Infrastructure/Site Reliability Engineer On-Call

Engg • San Francisco (CA)

On-site
USD 180,000 - 230,000
Senior SRE: AI-Driven Kubernetes Reliability at Scale
Senior SRE: AI-Driven Kubernetes Reliability at Scale

fal - Features & Labels • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+1
Senior SRE: AI Compute, Kubernetes & Observability
Senior SRE: AI Compute, Kubernetes & Observability

Justjoin • United States

Remote
USD 140,000 - 190,000
Health benefits
Financial planning
Family benefits
+2
Senior SRE: AI-Driven Ops & Incident Leader
Senior SRE: AI-Driven Ops & Incident Leader

Salesforce.com, inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 149,000 - 224,000
Senior SRE: Managed Kubernetes for AI Cloud Platforms
Senior SRE: Managed Kubernetes for AI Cloud Platforms

Lambda • Bellevue (WA)

On-site
USD 240,000 - 356,000
Health, dental, and vision coverage
401k matching
Flexible PTO
+1
Senior SRE: AI Infra, Hybrid Cloud & Automation
Senior SRE: AI Infra, Hybrid Cloud & Automation

d-Matrix • Santa Clara (CA)

On-site
USD 140,000 - 210,000
Senior Staff Cloud Reliability Engineer for AI Infra
Senior Staff Cloud Reliability Engineer for AI Infra

Epoch Biodesign • San Francisco (CA)

On-site
USD 180,000 - 220,000
Health insurance
401(k) with employer match
Paid Parental Leave
+2
Remote Senior SRE - AI Infra, Kubernetes & Terraform
Remote Senior SRE - AI Infra, Kubernetes & Terraform

Motion Recruitment • United States

Remote
USD 140,000 - 170,000
Medical, dental, and vision
Equity / Stock Options
Remote equipment stipend
+3
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
Senior DevOps Engineer - AI Infra, Kubernetes & Cloud
Senior DevOps Engineer - AI Infra, Kubernetes & Cloud

StackAI • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Hybrid work model
Remote work available
Office near Salesforce Park