Senior Software Engineer - Robinhood Command Center

Robinhood

California (MO)

On-site

USD 196,000 - 230,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
Equity ownership
401(k) matching
Lifestyle wallet
Life & disability insurance
Catered meals

Job summary

Robinhood is seeking a Senior Engineer for the newly formed Robinhood Command Center (RCC) in Menlo Park, CA. You will lead incident response, drive reliability and observability initiatives, and own tooling and processes for rapid, high-quality incident handling.

You will collaborate across product engineering, reliability, observability, and infrastructure teams to reduce customer impact and improve MTTD/MTTR.

Qualifications

  • 5+ years of software engineering experience and operating production systems.
  • 2+ years in reliability engineering, infrastructure, distributed systems, or production operations.
  • Hands-on incident leadership in roles like IMOC or incident commander.
  • Strong communication and cross-functional collaboration during high-severity incidents.
  • Deep knowledge of observability, fault-tolerance, and reliability patterns.

Responsibilities

  • Lead long-term reliability and observability strategy across Robinhood's infrastructure.
  • Collaborate with engineers to raise operational excellence and incident response.
  • Coordinate incident mitigation, rollbacks, and traffic shifts during incidents.
  • Develop incident management processes to minimize customer impact.
  • Own global dashboards and alerts tied to CUJs and business metrics.
  • Evolve incident response tooling, education, adoption, and MTTR/MTTD improvements.
  • Drive post-incident governance with postmortems and SEV reviews.
  • Design failure mitigation strategies to avoid full-region/datacenter failovers.
  • Build frameworks to improve monitoring and observability across services.
  • Own roadmap to bring observability to critical user journeys.
  • Deliver insights to executives about service quality and reliability.
  • Mentor and influence engineering culture and hiring.

Skills

Software engineering
Reliability engineering
Incident leadership
Cross-functional collaboration
Fault-tolerant design
Multi-region architectures

Tools

OpenTelemetry
Prometheus
Grafana

Job description

Join us in building the future of finance.

Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading.

About the team & role

We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards.

The Robinhood Command Center (RCC) is a newly formed reliability team that serves as the front line for detecting, coordinating, and mitigating production incidents across Robinhood.

As part of Robinhood’s broader reliability initiative, RCC works closely with product engineering, reliability, observability, infrastructure, and business teams to reduce customer impact and shorten incident duration.

As a Senior Engineer, you will be part of the founding RCC team, helping define how Robinhood responds to and learns from incidents at scale. This is a highly visible role focused on incident leadership, operational excellence, and reliability tooling. You will not own product services or core infrastructure, but you will own the processes and tools that enable fast, high-quality incident response.

This role is based in our Menlo Park, California office, with in-person attendance expected at least 3 days per week.

What you’ll do:
  • Serve as a senior technical leader driving the long-term reliability and observability strategy across Robinhood’s infrastructure
  • Partner closely across many different types of engineers to raise the bar for operational excellence and incident response
  • Lead incident mitigation efforts by coordinating service owners, facilitating time-sensitive decisions like rollbacks, traffic shifts, and maintaining a clear source of truth during active incidents
  • Develop and maintain incident management processes and procedures to ensure timely resolution and minimize customer impact
  • Own incident discovery at the company level by defining and maintaining global dashboards and alerts tied to critical user journeys (CUJs), availability, and business-impact metrics
  • Own and evolve incident response tooling and processes, including education, adoption, and measurement of MTTD/MTTR improvements
  • Drive post-incident governance and learning, defining standards for postmortems, SEV reviews, and follow-up tracking to ensure durable reliability improvements
  • Design and implement next-generation failure mitigation strategies that avoid full-region or full-datacenter failovers
  • Define and build frameworks to improve monitoring, alerting, and observability across hundreds of services and systems
  • Define and own the roadmap of bringing observability to critical user journeys for Robinhood’s products
  • Deliver key insights and executive-level reporting to enable better business decisions around service quality and reliability
  • Act as a force multiplier through mentoring, technical influence, and contributions to hiring and engineering culture
What you bring:
  • 5+ years of software engineering experience, including significant experience operating production systems
  • 2+ years focused on reliability engineering, infrastructure, distributed systems, or production operations
  • Hands-on experience serving in incident leadership roles (e.g., IMOC, incident commander, primary oncall)
  • Strong communication and cross-functional collaboration skills, especially during high-severity incidents
  • Deep knowledge of systems reliability, observability frameworks, and fault-tolerant architecture design
  • Experience with multi-region or multi-cluster architectures, capacity planning, and failover strategies
  • Familiarity with modern observability stacks (e.g., OpenTelemetry, Prometheus, Grafana)
  • Demonstrated ability to drive measurable improvements in MTTD, MTTR, availability, or customer impact
What we offer:
  • Challenging, high-impact work to grow your career
  • Performance driven compensation with multipliers for outsized impact, bonus programs, equity ownership, and 401(k) matching
  • Best in class benefits to fuel your work, including 100% paid health insurance for employees with 90% coverage for dependents
  • Lifestyle wallet - a highly flexible benefits spending account for wellness, learning, and more
  • Employer-paid life & disability insurance, fertility benefits, and mental health benefits
  • Time off to recharge including company holidays, paid time off, sick time, parental leave, and more!
  • Exceptional office experience with catered meals, events, and comfortable workspaces

In addition to the base pay range listed below, this role is also eligible for bonus opportunities + equity + benefits.

Base pay for the successful applicant will depend on a variety of job-related factors, which may include education, training, experience, location, business needs, or market demands. The expected base pay range for this role is based on the location where the work will be performed. For other locations not listed, compensation can be discussed with your recruiter during the interview process.

Base Pay Range:

Zone 1 (Menlo Park, CA; New York, NY; Bellevue, WA; Washington, DC)

$196,000-230,000 USD

Robinhood provides equal opportunity for all applicants, offers reasonable accommodations upon request, and complies with applicable equal employment and privacy laws. Inclusion is built into how we hire and work-welcoming different backgrounds, perspectives, and experiences so everyone can do their best. Please review the Privacy Policy for your country of application.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Robinhood Command Center
Senior Software Engineer, Robinhood Command Center

Robinhood • Bellevue (WA)

On-site
USD 196,000 - 230,000
Health insurance
Equity ownership
401(k) matching
+2
Senior Software Engineer - Robinhood Command Center
Senior Software Engineer - Robinhood Command Center

Robinhood • Menlo Park (CA)

On-site
USD 196,000 - 230,000
Performance-driven compensation
100% paid health insurance for employees
401(k) matching
Senior Software Engineer - Robinhood Command Center
Senior Software Engineer - Robinhood Command Center

Robinhood • Bellevue (WA)

On-site
USD 196,000 - 230,000
100% paid health insurance
Equity ownership
401(k) matching
+2
Staff Software Engineer, Observability
Staff Software Engineer, Observability

Robinhood • California (MO)

Hybrid
USD 180,000 - 270,000
Bonus opportunities
Equity ownership
Health insurance
+5
Staff Security Engineer, Detection & Response
Staff Security Engineer, Detection & Response

Robinhood • Bellevue (WA)

On-site
USD 217,000 - 255,000
Health insurance
Equity ownership
401(k) matching
+4
Senior Software Engineer, Load and Fault Environments
Senior Software Engineer, Load and Fault Environments

Socket.dev • Menlo Park (CA)

On-site
USD 196,000 - 230,000
Bonuses & equity
Health insurance
401(k) matching
+6
Senior Software Engineer, Load and Fault Environments
Senior Software Engineer, Load and Fault Environments

Robinhood • Menlo Park (CA)

On-site
USD 196,000 - 230,000
Health insurance
Equity ownership
401(k) matching
+2
Senior Staff Security Engineer
Senior Staff Security Engineer

Robinhood • Menlo Park (CA)

On-site
USD 247,000 - 290,000
Equity ownership
401(k) matching
Health insurance (employee)
Staff Security Engineer, Detection & Response
Staff Security Engineer, Detection & Response

Robinhood • Denver (CO)

Hybrid
USD 190,000 - 224,000
Health insurance
Equity and bonuses
401(k) matching
+2
Senior Software Engineer, Developer Experience (DevX)
Senior Software Engineer, Developer Experience (DevX)

Robinhood • Edison (CA)

Hybrid
USD 196,000 - 230,000
Health insurance
Equity ownership
401(k) matching
+2