Senior Reliability Engineer - Incident Command Center

Robinhood

Bellevue (WA)

On-site

USD 196,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

100% paid health insurance
Equity ownership
401(k) matching
Time off & holidays
Catered meals & office perks

Job summary

Robinhood is seeking a Senior Software Engineer for the Robinhood Command Center to lead reliability and observability across its infrastructure. You will collaborate with multiple engineering teams, drive incident response improvements, and own tooling and governance for post-incident learning.

This role emphasizes high-severity incident leadership, global dashboards, and end-to-end incident lifecycle governance to minimize customer impact and improve system availability.

Qualifications

  • 5+ years of software engineering experience, including significant experience operating production systems.
  • 2+ years focused on reliability engineering, infrastructure, distributed systems, or production operations.
  • Hands-on experience in incident leadership roles (e.g., IMOC, incident commander, primary oncall).
  • Strong communication and cross-functional collaboration skills during high-severity incidents.
  • Deep knowledge of systems reliability, observability frameworks, and fault-tolerant architecture design.
  • Experience with multi-region or multi-cluster architectures, capacity planning, and failover strategies.
  • Familiarity with modern observability stacks (OpenTelemetry, Prometheus, Grafana).
  • Demonstrated ability to drive measurable improvements in MTTD, MTTR, availability, or customer impact.

Responsibilities

  • Serve as a senior technical leader driving reliability and observability strategy across Robinhood’s infrastructure.
  • Partner across engineers to raise operational excellence and incident response.
  • Lead incident mitigation by coordinating service owners and managing rollbacks during active incidents.
  • Develop and maintain incident management processes to minimize customer impact.
  • Define global dashboards and alerts tied to critical user journeys and availability metrics.
  • Evolve incident response tooling and processes with education, adoption, and metrics.
  • Drive post-incident governance and postmortems to ensure durable improvements.
  • Design next-generation failure mitigation strategies to avoid full-region failovers.
  • Build frameworks to improve monitoring, alerting, and observability across services.
  • Own roadmap for bringing observability to critical user journeys.
  • Deliver insights and executive reporting to inform business decisions on service quality.
  • Mentor and influence engineering culture and hiring.

Skills

Incident leadership
Reliability engineering
Observability
Cross‑functional collaboration
Leadership
Multi-region architectures
OpenTelemetry
Prometheus
Grafana
MTTD/MTTR improvements

Tools

OpenTelemetry
Prometheus
Grafana

Job description

Robinhood is seeking a Senior Software Engineer for the Robinhood Command Center to lead reliability and observability across its infrastructure. You will collaborate with multiple engineering teams, drive incident response improvements, and own tooling and governance for post-incident learning.

This role emphasizes high-severity incident leadership, global dashboards, and end-to-end incident lifecycle governance to minimize customer impact and improve system availability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Reliability Engineer – Incident Command Leader
Senior Reliability Engineer – Incident Command Leader

Robinhood • Bellevue (WA)

On-site
USD 196,000 - 230,000
Health insurance
Equity ownership
401(k) matching
+2
Senior Reliability Engineer - Incident Command
Senior Reliability Engineer - Incident Command

Robinhood • New York (NY)

On-site
USD 120,000 - 150,000
Senior Reliability Engineer — Incident Leadership
Senior Reliability Engineer — Incident Leadership

Robinhood • Menlo Park (CA)

On-site
USD 196,000 - 230,000
Performance-driven compensation
100% paid health insurance for employees
401(k) matching
Senior Software Engineer, Robinhood Command Center
Senior Software Engineer, Robinhood Command Center

Robinhood • New York (NY)

On-site
USD 120,000 - 150,000
Senior Software Engineer, Robinhood Command Center
Senior Software Engineer, Robinhood Command Center

Robinhood • Bellevue (WA)

On-site
USD 196,000 - 230,000
Health insurance
Equity ownership
401(k) matching
+2
Senior Software Engineer — Resilience & Load Testing
Senior Software Engineer — Resilience & Load Testing

Robinhood • Menlo Park (CA)

On-site
USD 196,000 - 230,000
Health insurance
Equity ownership
401(k) matching
+2
Senior Software Engineer - Robinhood Command Center
Senior Software Engineer - Robinhood Command Center

Robinhood • Menlo Park (CA)

On-site
USD 196,000 - 230,000
Performance-driven compensation
100% paid health insurance for employees
401(k) matching
Senior Software Engineer, Load & Fault Resilience Platform
Senior Software Engineer, Load & Fault Resilience Platform

Robinhood • Bellevue (WA), Northern (KY)

Hybrid
USD 196,000 - 230,000
Health insurance
401(k) matching
Equity ownership
+2
Staff Security Engineer: AI-Driven Detection & Response
Staff Security Engineer: AI-Driven Detection & Response

Robinhood • Bellevue (WA)

On-site
USD 217,000 - 255,000
Health insurance
Equity ownership
401(k) matching
+4
Staff Observability Engineer - Scale, Telemetry & AI
Staff Observability Engineer - Scale, Telemetry & AI

Triwill Group • Menlo Park (CA)

Hybrid
USD 190,000 - 280,000
Health insurance
Equity ownership
401(k) matching
+1