Senior SRE Lead: Reliability & Cloud Platforms

Shield AI

San Mateo (CA)

On-site

USD 220,000 - 340,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Bonus
Benefits
Equity

Job summary

Shield AI is seeking an experienced SRE Lead to establish and mature a reliability program across our cloud infrastructure and platform services. You will define targets, improve observability, and guide production systems to operate and recover predictably.

This role is hands-on and technical, investigating complex failures, building tooling, and mentoring teams to embrace a reliability-first mindset while driving multi-quarter technical vision and roadmap.

Qualifications

  • 7+ years in SRE, software engineering, infrastructure engineering, or related fields.
  • Experience operating production services with defined availability and reliability requirements.
  • Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices.
  • Experience designing and operating infrastructure in AWS or another major cloud environment.
  • Experience with infrastructure-as-code and automated infrastructure provisioning.
  • Experience supporting containerized applications and distributed systems.
  • Experience developing operational tooling or automation using Python, Go, or a similar language.
  • Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services.
  • Experience leading incident response and root-cause analysis across engineering teams.
  • Experience leading and executing on technical vision of a team over multi-quarter timelines.

Responsibilities

  • Define and implement SLIs, SLOs, and other measures of service reliability
  • Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services
  • Lead technical response to complex incidents and drive root-cause analysis through resolution
  • Identify recurring failure modes and work with engineering teams to eliminate them
  • Improve system resilience through automation, testing, capacity planning, and failure recovery
  • Develop tooling and automation that reduces manual operational work
  • Partner with product and platform teams to incorporate reliability requirements into system design
  • Establish incident response practices that improve detection, diagnosis, communication, and recovery
  • Mentor product engineers and drive adoption of strong reliability and operational practices
  • Mentor teammates in SRE and Cloud Engineering
  • Define and manage short-and-long term SRE roadmap, distributing work across teammates

Skills

SRE experience
Incident response
Python/Go tooling
System observability

Tools

AWS
Kubernetes
Terraform

Job description

Shield AI is seeking an experienced SRE Lead to establish and mature a reliability program across our cloud infrastructure and platform services. You will define targets, improve observability, and guide production systems to operate and recover predictably.

This role is hands-on and technical, investigating complex failures, building tooling, and mentoring teams to embrace a reliability-first mindset while driving multi-quarter technical vision and roadmap.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE Lead: Reliability, Observability & Cloud
Senior SRE Lead: Reliability, Observability & Cloud

Shieldai • San Mateo (CA)

On-site
USD 180,000 - 240,000
SRE Lead: Reliability & Cloud Observability Architect
SRE Lead: Reliability & Cloud Observability Architect

BlackCube Labs • San Diego (CA)

On-site
USD 190,000 - 280,000
SRE Lead: Cloud Reliability & Incident Response
SRE Lead: Cloud Reliability & Incident Response

Shieldai • San Diego (CA)

On-site
USD 180,000 - 240,000
Sr. Staff Lead Site Reliability Engineer (R5803)
Sr. Staff Lead Site Reliability Engineer (R5803)

Shieldai • San Mateo (CA)

On-site
USD 180,000 - 240,000
Senior Production SRE: Cloud & On-Prem Reliability
Senior Production SRE: Cloud & On-Prem Reliability

Weights & Biases • New York (NY)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+15
Senior Staff Lead Site Reliability Engineer (R5803)
Senior Staff Lead Site Reliability Engineer (R5803)

Shieldai • San Diego (CA)

On-site
USD 180,000 - 240,000
Remote Senior Platform & SRE Lead
Remote Senior Platform & SRE Lead

United States Digital Space LLC • United States

Remote
USD 180,000 - 260,000
SRE Architecture Lead: Reliability & Cloud Platform
SRE Architecture Lead: Reliability & Cloud Platform

MACHINE LEARNING TECHNOLOGIES LLC • Atlanta (GA)

On-site
USD 140,000 - 190,000
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior SRE Leader: AI-Driven Reliability & Resilience
Senior SRE Leader: AI-Driven Reliability & Resilience

Jobtailor • Arlington (TX)

On-site
USD 180,000 - 240,000