Senior Infrastructure/Site Reliability Engineer On-Call

Engg

San Francisco (CA)

On-site

USD 180,000 - 230,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Engg, a San Francisco–based AI infrastructure company, seeks a senior on-call SRE to own incident response for Kubernetes, Ceph storage, and bare metal servers. You will coordinate with enterprise clients during outages, drive resolution timelines, and contribute to runbooks and automation.

You will lead cross-team investigations, manage etcd backups/restores, and help reduce incident frequency through reliability improvements and IaC tooling.

Qualifications

  • :years of production Kubernetes experience in enterprise environments.
  • Experience with distributed storage Ceph or equivalents.
  • Experience with at least one CNI plugin (Cilium/Calico).
  • Strong Linux admin and bare metal background.
  • Experience managing etcd backups/restores.

Responsibilities

  • Respond to and resolve production incidents across client infrastructure (Kubernetes, Ceph, bare metal).
  • Troubleshoot complex distributed systems issues during live outages.
  • Handle escalations requiring deep etcd, Ceph RGW, Cilium, and load balancer expertise.
  • Communicate incident status and timelines to enterprise clients.
  • Participate in follow-the-sun on-call rotation across time zones.
  • Document incidents and improve runbooks based on patterns.
  • Collaborate with infra on reliability and automation improvements.

Skills

Kubernetes operations
Distributed systems troubleshooting
On-call incident response
Linux administration
Client communication

Tools

Ceph
etcd
Cilium
NVIDIA Kubernetes Operator
Kubespray
Ansible

Job description

ABOUT THE ROLE

This is a senior on-call SRE role at an early-stage AI infrastructure company, where you will be the technical expert enterprise clients depend on when critical systems fail. You will own incident response across Kubernetes clusters, Ceph storage, and bare metal servers, keeping high-value AI workloads running at all times. Your calm judgment and deep distributed systems expertise will have a direct and immediate impact on client operations.

WHAT YOU'LL DO
  • Respond to and resolve production incidents across client infrastructure spanning Kubernetes, Ceph, and bare metal environments.
  • Troubleshoot complex distributed systems problems including pod scheduling failures, CNI networking issues, storage performance degradation, and hardware faults.
  • Handle escalations requiring deep expertise in etcd clusters, Ceph RGW authentication, Cilium networking, and bare metal load balancers.
  • Communicate directly with enterprise clients during incidents, providing clear status updates and resolution timelines.
  • Participate in a follow-the-sun on-call rotation with engineers across multiple time zones.
  • Document incidents thoroughly and improve runbooks based on recurring patterns.
  • Collaborate with the infrastructure team on long-term reliability improvements and automation to reduce incident frequency.
WHAT WE'RE LOOKING FOR
  • 5 or more years of production experience with Kubernetes in enterprise environments, including cluster operations, bare metal troubleshooting, admission controllers, and control plane architecture.
  • Deep production experience with distributed storage systems, particularly Ceph, or equivalent platforms such as Weka or VAST.
  • Production experience with at least one CNI plugin, preferably Cilium or Calico.
  • Strong modern Linux systems administration skills and comfort with bare metal infrastructure, IPMI, hardware troubleshooting, and networking.
  • Demonstrated ability to systematically diagnose and resolve complex distributed systems issues under pressure during live outages.
  • Production experience with etcd cluster management, including backup and restore procedures.
  • Experience with GPU infrastructure for AI/ML workloads, including the NVIDIA Kubernetes operator.
  • Familiarity with infrastructure-as-code tools such as Ansible, Kubespray, or similar orchestration frameworks.
  • Clear, composed communication with both technical and non-technical audiences during incidents.
  • Background in AI inference or training infrastructure is a strong plus.
LOCATION

On-site in San Francisco, CA. This role involves a follow-the-sun on-call rotation and requires timezone flexibility. Visa sponsorship is not available.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI Infra, On-Call, Kubernetes & Ceph
Senior SRE: AI Infra, On-Call, Kubernetes & Ceph

Engg • San Francisco (CA)

On-site
USD 180,000 - 230,000
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote

Motion Recruitment • United States

Remote
USD 140,000 - 170,000
Medical, dental, and vision
Equity / Stock Options
Remote equipment stipend
+3
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote
Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote

Motion Recruitment • Mount Laurel Township (NJ)

Remote
USD 130,000 - 180,000
Medical, dental, and vision benefits
Equity / Stock Options
Remote equipment stipend
+3
Member of Technical Staff, DevOps
Member of Technical Staff, DevOps

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Justjoin • United States

Remote
USD 140,000 - 190,000
Health benefits
Financial planning
Family benefits
+2
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

GCS Recruitment • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
AI Infrastructure Platform Operations Engineer remote in the US
AI Infrastructure Platform Operations Engineer remote in the US

Mirantis • United States

Remote
USD 110,000 - 150,000
Professional development
Conferences attendance
Team events
Software Engineer, Site Reliability
Software Engineer, Site Reliability

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health, dental, and vision insurance
Relocation assistance
Learning and growth opportunities
+1
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000