SRE Lead: Reliability & Cloud Observability Architect

BlackCube Labs

San Diego (CA)

On-site

USD 190,000 - 280,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Shield AI is seeking an experienced SRE Lead to drive the establishment and maturation of our reliability practices across cloud infrastructure and platform services. You will collaborate with Cloud Engineering and product teams to define targets, improve observability, and ensure production systems can be operated and recovered predictably.

This hands-on, deeply technical role involves investigating failures, improving tooling, and mentoring the team to adopt a reliability-first mindset across

Qualifications

  • 7+ years of experience in SRE or related fields.
  • Experience operating production services with defined availability and reliability requirements.
  • Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices.

Responsibilities

  • Define and implement SLIs, SLOs for service reliability.
  • Build monitoring, alerting, logging, and tracing for infrastructure.
  • Lead technical response to incidents and perform root-cause analysis.
  • Identify recurring failure modes and work with engineering teams to eliminate them.
  • Improve system resilience through automation, testing, capacity planning, and failure recovery.
  • Develop tooling and automation that reduces manual operational work.
  • Partner with product and platform teams to incorporate reliability requirements into system design.
  • Establish incident response practices that improve detection, diagnosis, communication, and recovery.
  • Mentor product engineers and drive adoption of strong reliability and operational practices.
  • Mentor teammates in SRE and Cloud Engineering.
  • Define and manage short-and-long term SRE roadmap, distributing work across teammates.

Skills

SRE leadership
Incident response
Root-cause analysis
Cloud reliability practices
Automation tooling
Distributed systems
Python/Go tooling
Technical vision
Troubleshooting

Tools

Kubernetes

Job description

Shield AI is seeking an experienced SRE Lead to drive the establishment and maturation of our reliability practices across cloud infrastructure and platform services. You will collaborate with Cloud Engineering and product teams to define targets, improve observability, and ensure production systems can be operated and recovered predictably.

This hands-on, deeply technical role involves investigating failures, improving tooling, and mentoring the team to adopt a reliability-first mindset across

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead: Reliability & Observability Champion (Cloud)
SRE Lead: Reliability & Observability Champion (Cloud)

Shield AI • San Diego (CA)

On-site
USD 190,000 - 280,000
Equity
Benefits
SRE Lead — Reliability, Observability & Incident Response
SRE Lead — Reliability, Observability & Incident Response

Shieldai • San Mateo (CA)

On-site
USD 180,000 - 260,000
Bonus
Benefits
Equity
Senior SRE Lead — Cloud Reliability & Platform Architect
Senior SRE Lead — Cloud Reliability & Platform Architect

Shieldai • San Diego (CA)

On-site
USD 150,000 - 230,000
Senior SRE Lead: Reliability, Automation & Incidents
Senior SRE Lead: Reliability, Automation & Incidents

Shield AI • San Diego (CA)

On-site
USD 183,000 - 275,000
Equity
Bonus
Benefits
Senior SRE Lead - Cloud Reliability & Observability
Senior SRE Lead - Cloud Reliability & Observability

Shield AI • San Mateo (CA)

On-site
USD 220,000 - 340,000
Equity
Bonus
Benefits
SRE Architecture Lead: Reliability & Cloud Platform
SRE Architecture Lead: Reliability & Cloud Platform

MACHINE LEARNING TECHNOLOGIES LLC • Atlanta (GA)

On-site
USD 140,000 - 190,000
Senior SRE Leader: Reliability & Observability
Senior SRE Leader: Reliability & Observability

Expedite Talent Solutions • United States

Hybrid
USD 130,000 - 160,000
SRE Manager: Lead Reliability & Observability at Scale
SRE Manager: Lead Reliability & Observability at Scale

Iac/interactivecorp • Sacramento (CA)

On-site
USD 150,000 - 210,000
Collaborative work environment
Commitment to carbon‑reduction mission
Flex schedule
+2
Senior Production SRE: Cloud & On-Prem Reliability
Senior Production SRE: Cloud & On-Prem Reliability

Weights & Biases • New York (NY)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+15
Senior SRE - Cloud & Observability
Senior SRE - Cloud & Observability

Ridgeline • Reno (NV)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Education reimbursement
Wellness reimbursement
+1