Reliability Engineer

OpenMind

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenMind in San Francisco is seeking a reliability engineer to architect testing, simulation, and observability infrastructure for autonomous systems. You will measure how often features fail and build automated reports that show reliability every morning.

You will work with Python/Go/C++ and robotics simulators, drive enterprise SLAs, and help define standards as we scale from lab prototypes to deployed fleets.

Qualifications

  • Experience in reliability engineering, test infrastructure, SRE, or simulation for autonomous systems
  • Strong engineering fundamentals in languages such as Python, Go, and/or C++
  • Experience with robotics simulation environments (Isaac Sim, Gazebo, MuJoCo, or proprietary equivalents)
  • A demonstrated obsession with measurement: you do not trust a system until you have quantified how it fails
  • The interpersonal skill to hold a high bar without becoming the person everyone avoids. Rigor with warmth is the job

Responsibilities

  • Build large-scale simulation pipelines that regression test every change against realistic deployed environments before release
  • Design and run structured reliability measurements: failure rates, mean time between failures, and degradation across hardware variants
  • Build fleet observability: telemetry ingestion, automated log collection, and AI-assisted scoring of robot performance at scale
  • Create the automated test harness that delivers a failure report on every engineer's desk each morning
  • Replay new software against our growing corpus of real-world deployment data to validate changes before they ship
  • Establish reliability and safety standards as we move into certifications, insurance requirements, and enterprise SLAs

Skills

Python
Go
C++

Tools

Isaac Sim
Gazebo
MuJoCo

Job description

The Role

The hardest problem in robotics is not the demo. It is the ten thousandth hour of operation. As OpenMind scales from lab prototypes to fleets deployed with enterprise customers, reliability becomes the product. You will build the testing, simulation, and observability infrastructure that tells us, every single morning, exactly how reliable our system is and what broke overnight. You will be the person who asks "how many times did this fail per 10,000 hours" about every new feature, and you will build the automated systems that answer that question before code ever reaches a robot. You will be our first reliability engineer, and you will define how we measure, test, and guarantee reliability as our fleet scales.

What You Will Do
  • Build large-scale simulation pipelines that regression test every change against realistic deployed environments before release
  • Design and run structured reliability measurements: failure rates, mean time between failures, and degradation across hardware variants
  • Build fleet observability: telemetry ingestion, automated log collection, and AI-assisted scoring of robot performance at scale
  • Create the automated test harness that delivers a failure report on every engineer's desk each morning
  • Replay new software against our growing corpus of real-world deployment data to validate changes before they ship
  • Establish reliability and safety standards as we move into certifications, insurance requirements, and enterprise SLAs
What We Look For
  • Experience in reliability engineering, test infrastructure, SRE, or simulation for autonomous systems
  • Strong engineering fundamentals in languages such as Python, Go, and/or C++
  • Experience with robotics simulation environments (Isaac Sim, Gazebo, MuJoCo, or proprietary equivalents)
  • A demonstrated obsession with measurement: you do not trust a system until you have quantified how it fails
  • The interpersonal skill to hold a high bar without becoming the person everyone avoids. Rigor with warmth is the job
Nice to Have
  • Background at an autonomy company (self-driving, drones, industrial robotics) with real deployment scale
  • Experience with hardware-in-the-loop testing or CI/CD for embedded systems
  • Familiarity with data infrastructure for large-scale log analysis
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Engineer — AI-Driven Robotics Fleet
Reliability Engineer — AI-Driven Robotics Fleet

OpenMind • San Francisco (CA)

On-site
USD 150,000 - 190,000
Staff Reliability Engineer, Hardware
Staff Reliability Engineer, Hardware

Bedrock Robotics • San Francisco (CA)

On-site
USD 120,000 - 150,000
Senior Reliability Engineer
Senior Reliability Engineer

Rhoda • Mountain View (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Requirements partnership
Industry background
Founding program experience
+3
Senior Reliability Engineer
Senior Reliability Engineer

Rhoda AI • Mountain View (CA)

On-site
USD 140,000 - 190,000
Senior Reliability Engineer
Senior Reliability Engineer

AeroVect Technologies Inc. • South San Francisco (CA)

On-site
USD 90,000 - 120,000
Robotics Infrastructure Engineer
Robotics Infrastructure Engineer

Tutor Intelligence • City of Watertown (NY)

On-site
USD 120,000 - 160,000
Robotics Infra & AI Automation Engineer
Robotics Infra & AI Automation Engineer

Tutor Intelligence • City of Watertown (NY)

On-site
USD 120,000 - 160,000
Robotics Hardware Reliability Engineer
Robotics Hardware Reliability Engineer

Dyna Robotics • Redwood City (CA)

On-site
USD 120,000 - 180,000
Fleet Response Engineer
Fleet Response Engineer

MVP Ventures • Mountain View (CA)

On-site
USD 170,000 - 230,000
Robotics Systems Integration Engineer
Robotics Systems Integration Engineer

Ultra • New York (NY)

On-site
USD 110,000 - 170,000