Senior Staff Lead Site Reliability Engineer

Shield AI

San Mateo (CA)

On-site

USD 150,000 - 230,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Excellent Medical Coverage
Stock Benefits
401K Matching
Flexible Work Hours
Onsite Gym

Job summary

Hivemind is seeking an experienced SRE Lead to establish and mature the reliability practices across our cloud infrastructure and platform services. You will collaborate with Cloud Engineering and product teams to define reliability targets, improve observability, and ensure production systems can be operated and recovered predictably.

This is a deeply technical, hands-on role focused on leading incident response, root-cause analysis, and building automation to reduce manual work while driving a

Qualifications

  • 7+ years of experience in SRE, software engineering, or related fields.

Responsibilities

  • Establish and mature the SRE function across cloud infrastructure and platform services.

Skills

SRE leadership
Incident response
Observability
Automation
Python/Go
Kubernetes
Cloud environments (AWS)
Infrastructure as Code
Monitoring and logging
Capacity planning

Tools

Kubernetes
AWS
Terraform
Python
Go

Job description

  • Hivemind is looking for an experienced SRE lead to help drive the establishment of our SRE function
  • As the SRE lead, you will establish and mature the reliability practices used across our cloud infrastructure and platform services
  • You will work with Cloud Engineering and product teams to define reliability targets, improve observability, and ensure that production systems can be operated and recovered predictably
  • This is a deeply technical, hands-on role
  • You will investigate complex failures, improve the systems and tooling used to operate our platforms, and turn lessons from incidents into engineering improvements
  • You will provide technical leadership for reliability engineering, helping teams adopt practices that improve system health without adding unnecessary process
  • You will act as a thought-leader and mentor within the Cloud Engineering and Reliability teams to level up teammates and encourage building with a reliability-first mindset
  • Define and implement SLIs, SLOs, and other measures of service reliability
  • Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services
  • Lead technical response to complex incidents and drive root-cause analysis through resolution
  • Identify recurring failure modes and work with engineering teams to eliminate them
  • Improve system resilience through automation, testing, capacity planning, and failure recovery
  • Develop tooling and automation that reduces manual operational work
  • Partner with product and platform teams to incorporate reliability requirements into system design
  • Establish incident response practices that improve detection, diagnosis, communication, and recovery
  • Mentor product engineers and drive adoption of strong reliability and operational practices
  • Mentor teammates in SRE and Cloud Engineering
  • Define and manage short-and-long term SRE roadmap, distributing work across teammates
Benefits
  • Excellent Medical Coverage
  • Mental Health Employee Assistance Program
  • Paid Parental Leave
  • Pet Insurance
  • Flexible Work Hours
  • Onsite Gym (DC)
  • Gym Discount (San Diego)
  • Free Parking
  • Competitive Compensation
  • Stock Benefits
  • 401K Services and Match
  • Experience leading and executing on technical vision of a team over multi-quarter timelines
  • Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services
  • Experience with infrastructure-as-code and automated infrastructure provisioning
  • Experience supporting containerized applications and distributed systems
  • Experience designing and operating infrastructure in AWS or another major cloud environment
  • 7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields
  • Experience leading incident response and root-cause analysis across engineering teams
  • Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices
  • Experience developing operational tooling or automation using Python, Go, or a similar language
  • Experience operating production services with defined availability and reliability requirements
  • Experience establishing or maturing an SRE function within an engineering organization
  • Experience with Kubernetes and cloud-native observability systems
  • Experience with capacity planning, performance analysis, and cloud cost management
  • Experience operating systems in regulated or compliance-driven environments
  • Background supporting shared infrastructure across multiple products or engineering organizations
  • Experience with building roadmaps in ticket systems
  • Experience acting as a mentor for other engineers
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Staff Site Reliability Engineer, SRE
Staff Site Reliability Engineer, SRE

Jobtailor • California (MO)

On-site
USD 120,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Staffing Science • Arizona

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

State of Wisconsin Investment Board • Madison (WI)

On-site
USD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
SRE Lead: Drive Reliability, Observability & Automation
SRE Lead: Drive Reliability, Observability & Automation

Shield AI • San Mateo (CA)

On-site
USD 150,000 - 230,000
Excellent Medical Coverage
Stock Benefits
401K Matching
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000