Staff Site Reliability Engineer (x/f/m)

Meyandy LLC

Berlin

Vor Ort

EUR 110.000 - 140.000

Vollzeit

Vor 12 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Meyandy LLC is seeking a Staff Site Reliability Engineer to strengthen platform reliability across our engineering teams in Berlin. You will lead reliability initiatives, mentor engineers, and drive cross-cutting improvements spanning infrastructure, observability, and incident management.

You will collaborate with software teams to embed reliability into development, define SLOs, and shape our on-call practices, with a focus on scalability and European-scale operations.

Qualifikationen

  • 8+ years in SRE/platform engineering in large-scale environments.
  • Proven containerization and Kubernetes deployment/scale experience.
  • Fluent in English; strong communication and collaboration skills.

Aufgaben

  • Lead large-scale cross-cutting reliability initiatives across the platform, spanning infrastructure automation, observability, and incident management.
  • Identify and drive improvements to incident detection, response, and postmortem analysis capabilities.
  • Define and evolve SLOs, error budgets, and alerting standards across multiple product teams.
  • Take part in the on-call rotation, and actively contribute to improving our on-call experience by refining alerting, reducing noise, and ensuring actionable telemetry.
  • Serve as a mentor and technical coach to senior engineers, helping elevate the craft of reliability engineering across the company.
  • Influence strategic decisions by providing technical guidance to leadership and representing reliability engineering in architectural reviews and platform discussions.
  • Partner with software engineering teams to embed reliability practices early in the development lifecycle.

Kenntnisse

SRE
Platform engineering
Kubernetes
Go
Python
Ruby
On-call experience
Mentoring
English

Tools

AWS
GCP
Azure
Observability tooling
Telemetry pipelines

Jobbeschreibung

Your Impact

We are looking for a Staff Site Reliability Engineer to join our SRE team dedicated to platform reliability within Platform Engineering.


Your mission will be to act as a technical leader driving Doctolib's reliability and scalability at a European scale, ensuring our platform remains reliable, debuggable, and resilient across infrastructure, observability, and cross-cutting reliability initiatives. You will play a pivotal role in a team driving reliability standards across 170+ applications, contributing directly to supporting 520,000 health professionals and 90 million patients in their daily healthcare journey.


This role sits at the intersection of infrastructure, developer experience, and product engineering. You'll act as a technical leader and strategic partner to SREs, software engineers, and product teams, guiding decisions, mentoring engineers, and driving cross-cutting initiatives that elevate our operational maturity.


What you'll do

Your responsibilities include but are not limited to:



  • Lead large-scale cross-cutting reliability initiatives across the platform, spanning infrastructure automation, observability, and incident management

  • Identify and drive improvements to incident detection, response, and postmortem analysis capabilities

  • Define and evolve SLOs, error budgets, and alerting standards across multiple product teams

  • Take part in the on-call rotation, and actively contribute to improving our on-call experience by refining alerting, reducing noise, and ensuring actionable telemetry

  • Serve as a mentor and technical coach to senior engineers, helping elevate the craft of reliability engineering across the company

  • Influence strategic decisions by providing technical guidance to leadership and representing reliability engineering in architectural reviews and platform discussions

  • Partner with software engineering teams to embed reliability practices early in the development lifecycle


Who you are

Before you read on: if you don't have the exact profile described below, but you feel this job description matches your skill set, we still encourage you to apply.


You'll be a great fit if you:



  • Have extensive experience (8+ years) in SRE, platform engineering, or infrastructure roles within a large-scale, multi-team production environment

  • Have proven experience with cloud platforms such as AWS, GCP, or Azure

  • Have strong experience with containerization and orchestration technologies, Kubernetes is a must, its deployment and scaling strategies ecosystem

  • Have implemented and operated SLIs, SLOs, and error budgets in production

  • Have experience managing on-call rotations and leading incident response in high-stakes environments

  • Have a strong systems engineering background with fluency in at least one backend programming language (e.g., Go, Python, Ruby)

  • Have a proven ability to lead through influence: setting technical direction, driving consensus, and mentoring engineers across teams

  • Are comfortable balancing long-term architecture work with fast, iterative improvements

  • Have clear, concise communication skills, both written and verbal, with the ability to drive alignment in ambiguous environments

  • Partner with feature teams to accelerate their production readiness, providing hands-on guidance on reliability best practices, launch reviews, and operational standards before go-live

  • Are fluent in English


It would be fantastic if you:



  • Have deep expertise in observability tooling and architecture (logging, tracing, metrics)

  • Have experience designing and operating high-scale telemetry pipelines and working with developers to improve instrumentation quality

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Staff Site Reliability Engineer (x/f/m)
Staff Site Reliability Engineer (x/f/m)

EngineersOfAI • Berlin

Hybrid
EUR 120.000 - 170.000
Senior Site Reliability Engineer (x/f/m)
Senior Site Reliability Engineer (x/f/m)

EngineersOfAI • Berlin

Vor Ort
EUR 90.000 - 140.000
Deutschlandticket
Engineering Manager - Site Reliability & Observability (x/f/m)
Engineering Manager - Site Reliability & Observability (x/f/m)

EngineersOfAI • Berlin

Vor Ort
EUR 120.000 - 150.000
Senior Site Reliability Engineer - Observability (x/f/m)
Senior Site Reliability Engineer - Observability (x/f/m)

Meyandy LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Staff Site Reliability Engineer (x/f/m)
Staff Site Reliability Engineer (x/f/m)

Doctolib • Berlin

Hybrid
EUR 110.000 - 150.000
Deutschlandticket (Germany-wide travel
28 vacation days
Hybrid work policy
+6
Director, Site Reliability Engineering
Director, Site Reliability Engineering

EngineersOfAI • Berlin

Vor Ort
EUR 90.000 - 130.000
Health insurance
Flexible working hours
Professional development opportunities
Senior Site Reliability Engineer (x/f/m)
Senior Site Reliability Engineer (x/f/m)

Doctolib • Berlin

Hybrid
EUR 90.000 - 140.000
Health insurance
Vacation days
Remote work flexibility
+1
Staff Site Reliability Engineer (x/f/m)
Staff Site Reliability Engineer (x/f/m)

DUDE CHEM • Berlin

Hybrid
EUR 120.000 - 160.000
Deutschlandticket
28 vacation days
Remote work days
+7
Engineering Manager - Observability & Reliability Engineering Obsession (x/f/m) Neu
Engineering Manager - Observability & Reliability Engineering Obsession (x/f/m) Neu

Doctolib GmbH • Berlin

Hybrid
EUR 120.000 - 170.000
Free comprehensive health insurance
Work from Berlin
Excellent team culture
Engineering Manager - Observability & Reliability Engineering Obsession (x/f/m)
Engineering Manager - Observability & Reliability Engineering Obsession (x/f/m)

Doctolib • Berlin

Vor Ort
EUR 120.000 - 160.000
Free comprehensive health insurance
ParentCare Program
Free mental health and coaching
+1