Staff Engineer

Talentify

Richardson (TX)

On-site

USD 150,000 - 190,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

GEICO is seeking an experienced Staff Engineer in the Technology Operations Center to lead incident management tooling and platform reliability. You will design, develop, and operate automation, dashboards, and data pipelines across incident response, on-call, and runbooks.

This role demands deep technical depth, strong leadership, and collaboration with SRE, platform, product, security, and business partners to raise reliability and reduce incidents.

Qualifications

  • Hands-on proficiency in Go, Java, Python, and C# for full-stack apps on Kubernetes and serverless platforms.
  • Experience with SQL and NoSQL databases and cloud-native services for incident data.
  • Experience building dashboards and data pipelines with Spark, Trino, Grafana, Superset, and Power BI.
  • Experience with OpenTelemetry and observability platforms like Grafana, Datadog, Splunk, Azure Monitor.
  • Experience with PagerDuty for incident management.
  • Proficiency in AI-assisted development tools such as Claude Code, Cursor, and GitHub Copilot.
  • Strong incident forensics and post-incident improvements; cross-team collaboration.
  • Solid software engineering fundamentals and large-scale production systems experience.

Responsibilities

  • Design, develop, and operate automation and data pipelines to scale incident management and on-call processes.
  • Build shared services and APIs to standardize incident response across teams.
  • Champion CI/CD, IaC, automated testing, rollback patterns, and production readiness.
  • Lead incidents as a technical leader, guiding troubleshooting and risk-based decisions.
  • Mentor engineers through reviews, documentation, and operational coaching.
  • Lead complex design reviews and align tooling with reliability goals.
  • Drive post-incident reviews and systemic reliability improvements.
  • Define incident response strategies and runbooks across integration points.

Skills

Go
Java
Python
C#

Tools

Kubernetes
Knative
Azure
AWS
Grafana
Datadog
Splunk
Azure Monitor
PagerDuty
Claude Code
Cursor
GitHub Copilot
Spark
Trino
Power BI

Job description

Why Join GEICO?

At GEICO, we offer a rewarding career where your ambitions are met with endless possibilities.

Every day we honor our iconic brand by offering quality coverage to millions of customers and being there when they need us most. We thrive on relentless innovation to exceed our customers' expectations while making a real impact on local communities nationwide.

Founded in 1936, GEICO is a member of the Berkshire Hathaway family of companies and one of the largest auto insurers in the United States. When you join our company, we want you to feel valued, supported, and proud to work here. That's why we offer the GEICO Pledge: Great Company, Great Culture, Great Rewards, and Great Careers.

Position Summary

Technology Operations Center is at the core of GEICO’s application and platform resiliency. It assists GEICO’s engineering teams with maintaining high availability of our customer and internally facing services while driving down time to detect and recover from incidents. It is governing our incident management processes and builds platforms that allow GEICO to manage and recover from incidents.

GEICO is seeking an experienced SRE Software Engineer with a passion for building, operating and troubleshooting high-performance, low-maintenance, zero-downtime complex distributed platforms and applications. You will help drive our transformation to a tech organization with engineering excellence and site reliability as its mission, by defining, implementing and operating our incident management processes and in‑house technology platforms that automate them.

This role focuses on improving Incident Management tooling and process across GEICO. It is a hands‑on technical leadership role focused on better incident management, faster time to detect, troubleshoot and recover from incidents and fewer repeat incidents, by designing, developing and operating the tools and processes that help all GEICO engineering teams manage their on‑call, runbooks, troubleshooting and BCDR.

Success in this role requires strong technical depth and process leadership. The right candidate can go deep on incident analysis, system behavior, and architecture, while also improving how teams run COEs, learn from incidents, and turn those lessons into engineering improvements.

Why This Role Is Different
  • This role blends deep technical understanding, hands‑on execution, and the ability to build software with real‑time incident leadership and platform and process improvements. You will help drive our incident management and response processes and the platforms that automate them.
  • You will shape how GEICO service engineering teams detect, manage, resolve, and learn from incidents.
  • Your work directly impacts the availability of GEICO’s critical applications and platforms and the experiences and satisfaction of millions of customers and tens of thousands of associates.
Position Responsibilities

As a Staff Engineer, you will:

Own and Evolve Enterprise‑Critical Platforms
  • Design, develop, and operate automation, self‑service tools, dashboards, and data pipelines that automate and scale our incident management, on‑call, paging, and troubleshooting processes.
  • Build shared services, APIs, data contracts, automation, and integrations that standardize incident response and reduce operational risk across teams.
  • Apply and help advance engineering standards across design, implementation, deployment, testing, observability, security, operational support, and production readiness.
  • Champion safe deployment, CI/CD, infrastructure as code, automated testing, rollback patterns, and operational controls that support frequent and reliable delivery.
  • Evaluate and implement modern technologies and tools that improve platform capability, compliance, visibility, reliability, and engineering effectiveness.
Serve as a Technical Leader During Incidents
  • Act as a technical leader and escalation point during high‑severity incidents, bringing sound architectural judgment, system‑level problem solving, and calm execution under pressure.
  • Guide troubleshooting strategy, cross‑team coordination, impact analysis, and risk‑based decision making to restore service safely and efficiently.
  • Lead or materially contribute to post‑incident reviews, root cause analysis, corrective action planning, and systemic reliability improvements.
  • Develop and maintain incident response strategies, operational runbooks, readiness criteria, triage models, and resilience practices across multiple integration points.
Influence Technical Direction and Engineering Culture
  • Lead complex design and architecture reviews spanning multiple teams, services, dependencies, and operational domains.
  • Partner with SRE, platform, product, infrastructure, security, and business stakeholders to align operational tooling with practical engineering needs and enterprise reliability goals.
  • Translate technical concepts, risks, and tradeoffs clearly for technical leaders, business stakeholders, and non‑technical audiences.
  • Mentor senior and mid‑level engineers through technical leadership, example, code and design reviews, documentation, and operational coaching.
  • Reinforce a culture of ownership, accountability, continuous improvement, psychological safety, learning, and operational excellence.

We have adopted a “You Build It, You Run It” strategy.

  • All our senior technologists take an active role in leading and managing high‑severity incidents requiring strong technical judgment, clear communication, and calm execution under pressure.
  • All our engineers have on‑call responsibilities as part of a 24x7 rotation supporting incident response and production support for mission‑critical platforms and processes they build and operate.
Qualifications
  • Hands‑on proficiency in multiple languages, including Go, Java, Python, and C#, for building production‑grade full‑stack applications on Kubernetes and serverless technologies such as Knative in Azure and AWS.
  • Experience with SQL and NoSQL technologies and cloud‑native services for storing and analyzing incident data.
  • Experience building and using data pipelines, analytics, and dashboards for operational metrics, trends, and KPIs using technologies such as Spark, Trino, Grafana, Superset, and Power BI.
  • Experience with OpenTelemetry and observability platforms such as Grafana, Datadog, Splunk, and Azure Monitor.
  • Experience with incident management platforms such as PagerDuty.
  • Proficiency with AI‑assisted development processes and tools such as Claude Code, Cursor, and GitHub Copilot.
  • Experience improving incident, post‑incident review, or reliability processes through automation, data, and cross‑team collaboration.
  • Strong incident forensics and root cause analysis skills, with the ability to improve COE quality through clear action items and follow‑through across teams.
  • Strong understanding of observability, reliability engineering, incident management, and post‑incident improvement practices.
  • Experience supporting incident response and high‑severity production incidents in complex environments.
  • Strong software engineering fundamentals and system design skills, with experience building reliable production systems at scale. 753
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Staff Engineer
Senior Staff Engineer

Government Employees Insurance Company • Bethesda (MD)

On-site
USD 140,000 - 210,000
Senior Engineer
Senior Engineer

Talentify • Richardson (TX)

On-site
USD 100,000 - 215,000
Senior Staff Engineer
Senior Staff Engineer

Talentify • Richardson (TX)

On-site
USD 120,000 - 260,000
Senior Staff Engineer - SRE - Incident Prevention / Post Incident Correction of Errors
Senior Staff Engineer - SRE - Incident Prevention / Post Incident Correction of Errors

Government Employees Insurance Company • Bethesda (MD)

On-site
USD 110,000 - 260,000
Senior Staff Engineer
Senior Staff Engineer

GEICO • Maryland

Hybrid
USD 120,000 - 260,000
Senior Staff Engineer SRE Incident Response (NOC)
Senior Staff Engineer SRE Incident Response (NOC)

GEICO • Chevy Chase (MD)

On-site
USD 110,000 - 260,000
401(k) plan with 6% match
Performance incentives
Tuition assistance
+2
Staff Engineer
Staff Engineer

Government Employees Insurance Company • Bethesda (MD)

On-site
USD 110,000 - 230,000
Senior Staff SRE Engineer — Incident Resilience Lead
Senior Staff SRE Engineer — Incident Resilience Lead

Talentify • Richardson (TX)

On-site
USD 120,000 - 260,000
Staff SRE Engineer — Incident Management & Reliability
Staff SRE Engineer — Incident Management & Reliability

Talentify • Richardson (TX)

On-site
USD 150,000 - 190,000
Senior Staff Engineer - SRE - Incident Prevention / Post Incident Correction of Errors
Senior Staff Engineer - SRE - Incident Prevention / Post Incident Correction of Errors

Geico • Seattle (WA)

On-site
USD 110,000 - 260,000