Senior SRE Engineer: Incident & Platform Reliability

Geico

Richardson (TX)

On-site

USD 100,000 - 215,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

GEICO is seeking an experienced Senior SRE Software Engineer to build and operate enterprise-grade platforms for incident management. You will design automation, dashboards, and data pipelines to support on-call, paging, and troubleshooting across high-availability services.

Lead incidents, perform post-incident reviews, and drive reliability improvements. Collaborate with SRE, platform, and security teams while contributing to engineering culture and operational excellence.

Qualifications

  • 4+ years of professional software engineering experience, preferably in platform engineering, reliability engineering, backend engineering, distributed systems, or operational tooling.
  • 3+ years of experience with architecture, design, system reliability, scalability, and technical delivery for production systems.
  • 2+ years of experience with open-source frameworks, modern engineering practices, or platform technologies.
  • 2+ years of experience with Azure, AWS, GCP, or another cloud service provider, or equivalent experience in complex hybrid environments.
  • Demonstrated ownership of production systems operating in 24x7 environments.
  • Strong software engineering fundamentals and system design skills, with experience building reliable production systems.

Responsibilities

  • Build and Operate Enterprise-Critical Platforms.
  • Design, develop and operate automation, self-service tools, dashboards, and data pipelines that automate and scale our incident management, on-call, paging and troubleshooting processes.
  • Build shared services, APIs, data contracts, automation, and integrations that standardize incident response and reduce operational risk.
  • Apply engineering standards across design, implementation, deployment, testing, observability, security, operational support, and production readiness.
  • Implement safe deployment, CI/CD, infrastructure as code, automated testing, rollback patterns, and operational controls that support frequent and reliable delivery.
  • Evaluate and implement modern technologies and tools that improve platform capability, compliance, visibility, reliability, and engineering effectiveness.
  • Serve as a Technical Leader During Incidents
  • Act as a technical leader during high-severity incidents, bringing sound technical judgment, system-level problem solving, and calm execution under pressure.
  • Contribute to troubleshooting strategy, cross-team coordination, impact analysis, and risk-based decision making to restore service safely and efficiently.
  • Contribute substantially to post-incident reviews, root cause analysis, corrective action planning, and systemic reliability improvements.
  • Develop and maintain operational runbooks, readiness criteria, triage models, and resilience practices across assigned integration points.
  • Contribute to Technical Direction and Engineering Culture
  • Contribute to design and architecture reviews spanning teams, services, dependencies, and operational domains.
  • Partner with SRE, platform, product, infrastructure, security, and business stakeholders to align operational tooling with practical engineering needs and enterprise reliability goals.
  • Translate technical concepts, risks, and tradeoffs clearly for technical leaders and non-technical stakeholders.
  • Mentor engineers through technical leadership, example, code and design reviews, documentation, and operational coaching.
  • Reinforce a culture of ownership, accountability, continuous improvement, psychological safety, learning, and operational excellence.

Skills

Go
Java
Python
C#

Education

Bachelor's degree in Computer Science or equivalent

Tools

Kubernetes
Azure
AWS
GCP
OpenTelemetry
Grafana
Datadog
Splunk
PagerDuty

Job description

GEICO is seeking an experienced Senior SRE Software Engineer to build and operate enterprise-grade platforms for incident management. You will design automation, dashboards, and data pipelines to support on-call, paging, and troubleshooting across high-availability services.

Lead incidents, perform post-incident reviews, and drive reliability improvements. Collaborate with SRE, platform, and security teams while contributing to engineering culture and operational excellence.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE Engineer: Incident Response & Reliability
Senior SRE Engineer: Incident Response & Reliability

Talentify • Richardson (TX)

On-site
USD 100,000 - 215,000
Senior Staff SRE Engineer — Incident & Reliability
Senior Staff SRE Engineer — Incident & Reliability

Geico • Richardson (TX)

On-site
USD 120,000 - 260,000
Senior Staff Engineer - Platform Reliability Lead
Senior Staff Engineer - Platform Reliability Lead

Government Employees Insurance Company • Bethesda (MD)

On-site
USD 140,000 - 210,000
Senior Staff SRE Engineer — Incident Resilience Lead
Senior Staff SRE Engineer — Incident Resilience Lead

Talentify • Richardson (TX)

On-site
USD 120,000 - 260,000
Staff SRE Engineer — Incident Management & Reliability
Staff SRE Engineer — Incident Management & Reliability

Talentify • Richardson (TX)

On-site
USD 150,000 - 190,000
Staff SRE Engineer: Incident & Reliability Lead
Staff SRE Engineer: Incident & Reliability Lead

Geico • Richardson (TX)

On-site
USD 110,000 - 230,000
Senior Staff Engineer, SRE & Incident Leadership
Senior Staff Engineer, SRE & Incident Leadership

GEICO • Maryland

Hybrid
USD 120,000 - 260,000
Senior Staff Engineer, SRE & Cloud Platform Lead
Senior Staff Engineer, SRE & Cloud Platform Lead

GEICO • Seattle (WA), Northern (KY)

Hybrid
USD 130,000 - 260,000
Staff Site Reliability Engineer - Incident Platform Lead
Staff Site Reliability Engineer - Incident Platform Lead

Government Employees Insurance Company • Bethesda (MD)

On-site
USD 110,000 - 230,000
Senior Staff SRE Engineer - COE & Incident Prevention
Senior Staff SRE Engineer - COE & Incident Prevention

Geico • Seattle (WA)

On-site
USD 110,000 - 260,000