SRE Lead: Cloud Reliability & Incident Response

Shieldai

San Diego (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Shield AI is seeking an experienced SRE Lead to establish and mature the reliability function across our cloud infrastructure and platform services. You will drive reliability targets, observability enhancements, and ensure production systems are operable and recoverable with a strong focus on automation and incident response.

You will mentor engineers and shape the SRE roadmap, partnering with Cloud Engineering and product teams to embed reliability into system design and outages prevention

Qualifications

  • 7+ years of experience in SRE, software or infrastructure engineering
  • Experience operating production services with defined availability requirements
  • Experience implementing SLIs, SLOs, monitoring and incident response practices
  • Experience designing and operating infrastructure in AWS or another major cloud
  • Experience with infrastructure-as-code and automated provisioning
  • Experience supporting containerized applications and distributed systems
  • Experience developing operational tooling or automation using Python or Go
  • Ability to diagnose complex failures across apps, infra, networking and services
  • Experience leading incident response and root-cause analyses
  • Experience guiding multi-quarter technical roadmaps

Responsibilities

  • Define and implement SLIs, SLOs and reliability metrics
  • Build and improve monitoring, logging, tracing and alerting for infra and platform services
  • Lead incident response and drive root-cause analysis to resolution
  • Identify recurring failure modes and eliminate them with engineering teams
  • Improve system resilience via automation, testing, capacity planning
  • Develop tooling to reduce manual operational work
  • Incorporate reliability requirements into system design with product teams
  • Mentor engineers and promote reliability-first practices
  • Define and manage SRE roadmap across the team

Skills

SRE Leadership
Cloud engineering
Incident response
Python/Go
Observability
Kubernetes
AWS
Infrastructure as code

Tools

Kubernetes
AWS
Terraform
Python
Go

Job description

Shield AI is seeking an experienced SRE Lead to establish and mature the reliability function across our cloud infrastructure and platform services. You will drive reliability targets, observability enhancements, and ensure production systems are operable and recoverable with a strong focus on automation and incident response.

You will mentor engineers and shape the SRE roadmap, partnering with Cloud Engineering and product teams to embed reliability into system design and outages prevention

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead: Reliability & Cloud Observability Architect
SRE Lead: Reliability & Cloud Observability Architect

BlackCube Labs • San Diego (CA)

On-site
USD 190,000 - 280,000
Senior SRE Lead: Reliability, Observability & Cloud
Senior SRE Lead: Reliability, Observability & Cloud

Shieldai • San Mateo (CA)

On-site
USD 180,000 - 240,000
Senior SRE Lead: Reliability & Cloud Platforms
Senior SRE Lead: Reliability & Cloud Platforms

Shield AI • San Mateo (CA)

On-site
USD 220,000 - 340,000
Bonus
Benefits
Equity
SRE Lead: Drive Reliability, Observability & Automation
SRE Lead: Drive Reliability, Observability & Automation

Shield AI • San Mateo (CA)

On-site
USD 150,000 - 230,000
Excellent Medical Coverage
Stock Benefits
401K Matching
+2
SRE Architecture Lead: Reliability & Cloud Platform
SRE Architecture Lead: Reliability & Cloud Platform

MACHINE LEARNING TECHNOLOGIES LLC • Atlanta (GA)

On-site
USD 140,000 - 190,000
SRE Lead: Incident Commander & Reliability Champion
SRE Lead: Incident Commander & Reliability Champion

U.S. Bank • Northern (KY)

Hybrid
USD 112,000 - 131,000
Healthcare
Retirement plan
Paid vacation
+2
Product SRE Lead — AI-Driven Reliability & Cloud
Product SRE Lead — AI-Driven Reliability & Cloud

Cvent, Inc. • Tysons (VA)

On-site
USD 180,000 - 240,000
SRE Lead: Incident Commander for Cloud Reliability
SRE Lead: Incident Commander for Cloud Reliability

Us Bank • Atlanta (GA)

On-site
USD 112,000 - 131,000
Healthcare
Life Insurance
Disability
+6
Senior Production SRE: Cloud & On-Prem Reliability
Senior Production SRE: Cloud & On-Prem Reliability

Weights & Biases • New York (NY)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+15
Senior SRE: Cloud Reliability & Incidents Lead
Senior SRE: Cloud Reliability & Incidents Lead

Illumio • San Jose (CA)

On-site
USD 120,000 - 150,000