Senior Site Reliability Engineer, Vehicle SW

Icehouseventures

Germany (OH)

Hybrid

USD 79,944 - 125,627

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Wayve in Germany, Baden-Württemberg, is seeking a Site Reliability Engineer to keep the autonomous vehicle software fleet reliable, observable, and safe on public roads. You will work at the boundary of software, hardware, and operations, turning real-world incidents into lasting engineering improvements.

We are looking for strong SRE experience, Linux fundamentals, and expertise with CI/CD, containers, and orchestration.

Qualifications

  • Proven experience in SRE or platform operations for distributed systems.
  • Strong Linux fundamentals with CI/CD, Docker, and Kubernetes.
  • Proficiency in Python, C++, or Rust with automation focus.
  • Troubleshoot networking, distributed systems, and DB performance issues.
  • Experience designing observability stacks with Datadog, Prometheus, Grafana, OpenTelemetry, Splunk or Humio.
  • Clear communication and incident leadership, post-mortems.

Responsibilities

  • Own reliability, availability, and performance of vehicle software systems.
  • Participate in on-call rotation and provide off-hours support.
  • Build and operate monitoring, logging, alerting, and on-call tooling.
  • Lead incident response and post-incident learning with durable fixes.
  • Design automation for fleet ops, deployments, and workflows.
  • Collaborate with Vehicle SW, operations, and platform teams on SLOs.
  • Hardening production via capacity planning and change management.

Skills

SRE/Platform ops
Linux fundamentals
CI/CD, Docker, Kubernetes
Scripting: Python/C/C++/Rust
Observability stack
Incident leadership
Communication

Tools

Datadog
Prometheus
Grafana
OpenTelemetry
Splunk
Humio

Job description

The role

As an SRE in Vehicle Software, you will keep Wayve’s autonomous driving fleet reliable, observable, and safe while it operates on public roads. You will work at the boundary of software, hardware, and operations, turning real-world incidents and performance bottlenecks into lasting engineering improvements. This role offers a direct line of sight from what you build to safer deployments, faster iteration, and greater fleet scale.

Key responsibilities
  • Own and improve the reliability, availability, and performance of vehicle software systems used across the dev fleet.
  • Take part in a team on‑call rotation, providing out‑of‑hours support for live systems when required.
  • Build and operate monitoring, logging, alerting, and on‑call tooling that enables fast detection, diagnosis, and recovery.
  • Drive incident response and post‑incident learning, translating root causes into durable fixes and preventive controls.
  • Design and deliver automation for fleet operations, deployments, and repetitive workflows to reduce manual intervention.
  • Partner closely with Vehicle SW, operations, and platform teams to define SLOs, reliability metrics, and release readiness.
  • Continuously harden the production environment through capacity planning, change management, and reliability‑focused reviews.
About you

In order to set you up for success as a Site Reliability Engineer at Wayve, we’re looking for the following skills and experience.

Essential skills
  • Proven experience in an SRE, production reliability, or platform operations role for complex distributed systems.
  • Strong Linux fundamentals and hands‑on experience with CI/CD, containers (Docker), and orchestration (Kubernetes).
  • Proficiency in at least one systems or scripting language (Python, C++, or Rust) with a bias for automation.
  • Deep troubleshooting skills across networking, distributed systems, and databases, including performance and availability issues.
  • Experience designing observability stacks and using tools such as Datadog, Prometheus, Grafana, OpenTelemetry, Splunk, or Humio.
  • Clear communication skills, including incident leadership, writing post‑mortems, and influencing engineering priorities.
Desirable skills
  • Cloud platform experience (AWS, GCP, or Azure), including infrastructure‑as‑code and secure production operations.
  • Experience with real‑time or safety‑critical systems, hardware‑in‑the‑loop, or embedded/robotics environments.
  • Familiarity with fleet operations, telemetry pipelines, and operating software on edge devices at scale.
  • Experience defining and running SLOs/SLIs and reliability programs across multiple teams.

This is a full‑time role based in our office in Germany, Baden-Württemberg (Hybrid 3 days a week min). At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home.

Wayve is committed to creating an inclusive interview experience. If you require any accommodations or adjustments to participate fully in our interview process, please let us know.

We understand that everyone has a unique set of skills and experiences and that not everyone will meet all of the requirements listed above. If you’re passionate about self‑driving cars and think you have what it takes to make a positive impact on the world, we encourage you to apply.

At Wayve we’re committed to creating a diverse, fair and respectful culture that is inclusive of everyone based on their unique skills and perspectives, and regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, veteran status, pregnancy or related condition (including breastfeeding) or any other basis as protected by applicable law.

DISCLAIMER: We will not ask about marriage or pregnancy, care responsibilities or disabilities in any of our job adverts or interviews. However, we do look to capture information about care responsibilities, and disabilities among other diversity information as part of an optional DEI Monitoring form to help us identify areas of improvement in our hiring process and ensure that the process is inclusive and non‑discriminatory.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer, Vehicle SW
Senior Site Reliability Engineer, Vehicle SW

EngineersOfAI • Germany (OH)

On-site
USD 90,000 - 130,000
Site Reliability Engineering Manager, Vehicle Software
Site Reliability Engineering Manager, Vehicle Software

Wayve • Sunnyvale (CA)

Hybrid
USD 276,000 - 311,000
Hybrid work policy
Equity package
Senior Software Engineer, Fleet Management
Senior Software Engineer, Fleet Management

Socket.dev • Sunnyvale (CA)

Hybrid
USD 210,000 - 267,000
Application Software Engineer
Application Software Engineer

Icehouseventures • Germany (OH)

On-site
USD 81,000 - 105,000
Senior Software Engineer, Fleet Management
Senior Software Engineer, Fleet Management

Wayve • Sunnyvale (CA)

Hybrid
USD 210,000 - 267,000
Competitive equity package
Hybrid work policy
Senior Software Engineer, Fleet Management
Senior Software Engineer, Fleet Management

Necessary Ventures • Sunnyvale (CA)

On-site
USD 210,000 - 267,000
Software Engineer, Fleet Management
Software Engineer, Fleet Management

Worky • Sunnyvale (CA)

Hybrid
USD 177,000 - 221,000
Senior Field Engineer
Senior Field Engineer

Wayve • Detroit (MI)

On-site
USD 120,000 - 160,000
Software Engineer, Fleet Management
Software Engineer, Fleet Management

Wayve • Sunnyvale (CA)

Hybrid
USD 177,000 - 221,000
Safety System Engineer
Safety System Engineer

Icehouseventures • Sunnyvale (CA)

Hybrid
USD 100,000 - 150,000