Senior Reliability Engineer – Incident Command Leader

Robinhood

Bellevue (WA)

On-site

USD 196,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Equity ownership
401(k) matching
Bonus programs
Paid time off

Job summary

Robinhood is seeking a Senior Software Engineer to join the Command Center in New York. You will lead reliability and observability initiatives across Robinhood’s infrastructure, coordinating with multiple engineering teams to improve incident response and service quality.

The role focuses on developing tooling, dashboards, and governance to minimize customer impact. This high-visibility position offers opportunities for mentoring, substantial influence on engineering practices, and a

Qualifications

  • 5+ years of software engineering experience, including significant experience operating production systems.
  • 2+ years focused on reliability engineering, infrastructure, distributed systems, or production operations.
  • Hands‑on experience serving in incident leadership roles (e.g., IMOC, incident commander, primary on‑call).
  • Strong communication and cross‑functional collaboration skills, especially during high‑severity incidents.
  • Deep knowledge of systems reliability, observability frameworks, and fault‑tolerant architecture design.
  • Experience with multi‑region or multi‑cluster architectures, capacity planning, and failover strategies.
  • Familiarity with modern observability stacks (e.g., OpenTelemetry, Prometheus, Grafana).
  • Demonstrated ability to drive measurable improvements in MTTD, MTTR, availability, or customer impact.

Responsibilities

  • Serve as a senior technical leader driving the long-term reliability and observability strategy across Robinhood’s infrastructure
  • Partner closely across many different types of engineers to raise the bar for operational excellence and incident response
  • Lead incident mitigation efforts by coordinating service owners, facilitating time-sensitive decisions like rollbacks, traffic shifts, and maintaining a clear source of truth during active incidents
  • Develop and maintain incident management processes and procedures to ensure timely resolution and minimize customer impact
  • Own incident discovery at the company level by defining and maintaining global dashboards and alerts tied to critical user journeys (CUJs), availability, and business-impact metrics
  • Own and evolve incident response tooling and processes, including education, adoption, and measurement of MTTD/MTTR improvements
  • Drive post-incident governance and learning, defining standards for postmortems, SEV reviews, and follow-up tracking to ensure durable reliability improvements
  • Design and implement next-generation failure mitigation strategies that avoid full-region or full-datacenter failovers
  • Define and build frameworks to improve monitoring, alerting, and observability across hundreds of services and systems
  • Define and own the roadmap of bringing observability to critical user journeys for Robinhood’s products
  • Deliver key insights and executive-level reporting to enable better business decisions around service quality and reliability
  • Act as a force multiplier through mentoring, technical influence, and contributions to hiring and engineering culture

Skills

Software engineering
Reliability engineering
Incident leadership
Communication
Systems reliability
Observability
Multi-region architectures
Capacity planning
Failover strategies
OpenTelemetry
Prometheus
Grafana
MTTD/MTTR improvements

Tools

OpenTelemetry
Prometheus
Grafana

Job description

Robinhood is seeking a Senior Software Engineer to join the Command Center in New York. You will lead reliability and observability initiatives across Robinhood’s infrastructure, coordinating with multiple engineering teams to improve incident response and service quality.

The role focuses on developing tooling, dashboards, and governance to minimize customer impact. This high-visibility position offers opportunities for mentoring, substantial influence on engineering practices, and a

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Reliability Engineer - Incident Command Center
Senior Reliability Engineer - Incident Command Center

Robinhood • Bellevue (WA)

On-site
USD 196,000 - 230,000
100% paid health insurance
Equity ownership
401(k) matching
+2
Senior Reliability Engineer - Incident Command
Senior Reliability Engineer - Incident Command

Robinhood • New York (NY)

On-site
USD 120,000 - 150,000
Senior Reliability Engineer – Incident Leadership
Senior Reliability Engineer – Incident Leadership

Robinhood • Edison (CA)

Hybrid
USD 196,000 - 230,000
Performance-driven compensation
401(k) matching
100% paid health insurance for employees
Senior Reliability Engineer — Incident Leadership
Senior Reliability Engineer — Incident Leadership

Robinhood • Menlo Park (CA)

On-site
USD 196,000 - 230,000
Performance-driven compensation
100% paid health insurance for employees
401(k) matching
Senior Software Engineer, Robinhood Command Center
Senior Software Engineer, Robinhood Command Center

Robinhood • New York (NY)

On-site
USD 120,000 - 150,000
Senior Software Engineer, Robinhood Command Center
Senior Software Engineer, Robinhood Command Center

Robinhood • Bellevue (WA)

On-site
USD 196,000 - 230,000
Health insurance
Equity ownership
401(k) matching
+2
Senior Software Engineer - Robinhood Command Center
Senior Software Engineer - Robinhood Command Center

Robinhood • Edison (CA)

Hybrid
USD 196,000 - 230,000
Performance-driven compensation
401(k) matching
100% paid health insurance for employees
Senior Software Engineer - Robinhood Command Center
Senior Software Engineer - Robinhood Command Center

Robinhood • Menlo Park (CA)

On-site
USD 196,000 - 230,000
Performance-driven compensation
100% paid health insurance for employees
401(k) matching
Senior Backend Engineer - Real-Time Distributed Systems
Senior Backend Engineer - Real-Time Distributed Systems

Robinhood • New York (NY), Menlo Park (CA)

Hybrid
USD 196,000 - 230,000
Competitive benefits package
401(k) matching
Equity ownership
+1
Staff Observability Engineer — Lead Telemetry at Scale
Staff Observability Engineer — Lead Telemetry at Scale

Robinhood • Menlo Park (CA)

On-site
USD 210,000 - 290,000
Equity ownership
401(k) matching
Health insurance for employees