Senior SRE Lead - Cloud Reliability & Observability

Shield AI

San Mateo (CA)

On-site

USD 220,000 - 340,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Bonus
Benefits

Job summary

Shield AI is a venture-backed defense-tech company with the mission of protecting service members and civilians with intelligent systems. As the SRE lead, you will establish and mature the reliability practices used across our cloud infrastructure and platform services.

You will work with Cloud Engineering and product teams to define reliability targets, improve observability, and ensure that production systems can be operated and recovered predictably. This is a deeply technical, hands-on role.

Qualifications

  • 7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields.
  • Experience operating production services with defined availability and reliability requirements.
  • Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices.
  • Experience designing and operating infrastructure in AWS or another major cloud environment.
  • Experience with infrastructure-as-code and automated infrastructure provisioning.
  • Experience supporting containerized applications and distributed systems.
  • Experience developing operational tooling or automation using Python, Go, or a similar language.
  • Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services.
  • Experience leading incident response and root-cause analysis across engineering teams.
  • Experience leading and executing on technical vision of a team over multi-quarter timelines.

Responsibilities

  • Define and implement SLIs, SLOs, and other measures of service reliability.
  • Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services.
  • Lead technical response to complex incidents and drive root-cause analysis through resolution.
  • Identify recurring failure modes and work with engineering teams to eliminate them.
  • Improve system resilience through automation, testing, capacity planning, and failure recovery.
  • Develop tooling and automation that reduces manual operational work.
  • Partner with product and platform teams to incorporate reliability requirements into system design.
  • Establish incident response practices that improve detection, diagnosis, communication, and recovery.
  • Mentor product engineers and drive adoption of strong reliability and operational practices.
  • Mentor teammates in SRE and Cloud Engineering.
  • Define and manage short-and-long term SRE roadmap, distributing work across teammates.

Skills

SRE/DevOps
Cloud engineering
Incident response
Python/Go
Automation tooling

Tools

AWS
Kubernetes
Terraform

Job description

Shield AI is a venture-backed defense-tech company with the mission of protecting service members and civilians with intelligent systems. As the SRE lead, you will establish and mature the reliability practices used across our cloud infrastructure and platform services.

You will work with Cloud Engineering and product teams to define reliability targets, improve observability, and ensure that production systems can be operated and recovered predictably. This is a deeply technical, hands-on role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead: Reliability & Observability Champion (Cloud)
SRE Lead: Reliability & Observability Champion (Cloud)

Shield AI • San Diego (CA)

On-site
USD 190,000 - 280,000
Equity
Benefits
Senior SRE Lead — Cloud Reliability & Platform Architect
Senior SRE Lead — Cloud Reliability & Platform Architect

Shieldai • San Diego (CA)

On-site
USD 150,000 - 230,000
SRE Lead: Reliability & Cloud Observability Architect
SRE Lead: Reliability & Cloud Observability Architect

BlackCube Labs • San Diego (CA)

On-site
USD 190,000 - 280,000
Senior SRE Lead: Reliability, Automation & Incidents
Senior SRE Lead: Reliability, Automation & Incidents

Shield AI • San Diego (CA)

On-site
USD 183,000 - 275,000
Equity
Bonus
Benefits
SRE Lead — Reliability, Observability & Incident Response
SRE Lead — Reliability, Observability & Incident Response

Shieldai • San Mateo (CA)

On-site
USD 180,000 - 260,000
Bonus
Benefits
Equity
Sr. Staff Site Reliability Engineer (R5803)
Sr. Staff Site Reliability Engineer (R5803)

Shieldai • San Mateo (CA)

On-site
USD 180,000 - 260,000
Bonus
Benefits
Equity
Sr. Staff Site Reliability Engineer (R5803)
Sr. Staff Site Reliability Engineer (R5803)

Shieldai • San Diego (CA)

On-site
USD 150,000 - 230,000
Sr. Staff Lead Site Reliability Engineer (R5803)
Sr. Staff Lead Site Reliability Engineer (R5803)

Shield AI • San Mateo (CA)

On-site
USD 220,000 - 340,000
Equity
Bonus
Benefits
Senior Staff Lead Site Reliability Engineer (R5803)
Senior Staff Lead Site Reliability Engineer (R5803)

BlackCube Labs • San Diego (CA)

On-site
USD 190,000 - 280,000
Sr. Staff Lead Site Reliability Engineer (R5803)
Sr. Staff Lead Site Reliability Engineer (R5803)

Shield AI • San Diego (CA)

On-site
USD 190,000 - 280,000
Equity
Benefits