Site Reliability Engineer - Datadog, Kafka (d/f/m)

Personio

München

Vor Ort

EUR 70.000 - 90.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

28 days of paid vacation
Mental health support
Competitive reward package
Pre-IPO equity
Flex Days

Zusammenfassung

Personio is seeking a highly skilled Site Reliability Engineer (SRE) in Munich. This role requires 2 days a week in the office and involves engaging in the full service lifecycle, monitoring live services, and ensuring sustainable scalability through automation. Ideal candidates will have 6+ years in SaaS software development, particularly in distributed systems, and possess strong skills in Java, Kotlin, and observability tools. The position offers a competitive compensation package, 28 days of vacation, and benefits supporting family and mental health.

Qualifikationen

  • 6+ years of experience with SaaS software development in distributed systems.
  • 2+ years’ experience as an SRE or similar role.
  • Hands-on experience running Kafka at scale.

Aufgaben

  • Engage in and improve the full service lifecycle from design to deployment.
  • Operate, monitor, and maintain live services.
  • Support incident management processes, including design and metrics.

Kenntnisse

SaaS software development
Distributed systems
Kotlin
Java
Typescript
Python
Docker
Kubernetes
Datadog
Kafka

Ausbildung

Bachelor’s degree in Computer Science or equivalent

Tools

CI/CD tooling

Jobbeschreibung

Personio's intelligent HR platform helps small and medium-sized organizations unlock the power of people by making complicated, time‑consuming tasks simple and efficient. Our team of 1,500 Personios is building user-friendly products that delight our 15,000+ customers and their 1.5 million employees. Ready to make an impact from day one?

The Role

This role requires 2 days a week in our Munich, Berlin or Dublin office. Join us to shape the future of software in the underserved and high‑impact HR technology industry. Your work will have a direct and tangible impact on customers, offering ownership and the chance to make a meaningful difference. As we prepare for significant growth, you’ll face exciting challenges and have the opportunity to influence our path toward becoming one of the world’s leading tech companies.

What You’ll Do
  • Engage in and improve the full service lifecycle from initial design through deployment, operation, and continuous improvement.
  • Prepare services for production by taking part in system design reviews, developing shared frameworks and platforms, planning capacity and conducting launch assessments.
  • Operate, monitor, and maintain live services, designing observability stacks and dashboards to track key metrics and improve operational insight.
  • Ensure sustainable scalability through automation, actively contributing to continuous improvement for reliability and delivery speed.
  • Collaborate with product and engineering teams to define SLOs, error budgets and ensure services are reliable, scalable and observable.
  • Support incident management processes, including on‑call rotations, assisting with outage response, and contributing to post‑mortems and root cause analysis.
  • Identify and reduce toil through process automation, creating playbooks and automated runbooks to reduce MTTR.
  • Support resilience strategies and help implement chaos testing to proactively uncover weaknesses and validate recovery strategies.
  • Own and maintain the reliability of our event streaming and Change Data Capture (CDC) stack.
  • Mentor and train peers on reliability best practices and tooling, contributing to community growth.
What You Need To Succeed
  • Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
  • 6+ years of experience with SaaS software development in distributed systems using languages such as Kotlin/Java, Typescript, Python, and technologies like IaC, Docker, and Kubernetes.
  • 2+ years’ experience as an SRE or similar role designing, operating, analyzing and troubleshooting distributed systems in agile environments.
  • Act as a Datadog subject matter expert, assisting with observability stack design, dashboard creation, and training peers in best practices.
  • Hands‑on experience running Kafka at scale including configuration, operational failure modes and reliable recovery/runbooks.
  • Systematic problem solving and debugging skills with a strong sense of ownership and bias towards establishing mechanisms which can scale across the entire company.
  • Excellent written, verbal, and documentation skills.
  • Collaborative team player, able to communicate effectively across disciplines.
Nice to Have / Bonus
  • Experience with CI/CD tooling (GitHub Actions/GitOps tools)
  • Experience tuning JVM‑based services and Node.js runtimes
  • Experience with AWS MSK Connect
Why Personio

Personio is an equal opportunities employer, committed to building an integrative culture where everyone feels welcomed and supported. We embrace uniqueness and understand that our diverse, values‑driven culture makes us stronger. We are proud to have an inclusive workplace environment that will foster your development no matter your gender, civil status, family status, sexual orientation, religion, age, disability, education level, or race.

At Personio, we value in‑person collaboration while also offering flexibility. This role is office‑based, with 2 required in your contracted office location. The remaining days can be worked from home or in the office if you prefer. In addition, you’ll have 20 Flex Days per year to work remotely from other locations.

Aside from our people, culture and mission, check out some of the other benefits that make Personio a great place to work:

  • Receive a competitive reward package – reevaluated each year – that includes salary, benefits, and pre‑IPO equity.
  • Enjoy 28 days of paid vacation, plus an additional day after 2 and 4 years.
  • Make an impact on the environment and society with 1 (fully paid) Impact Day.
  • Receive generous family leave, child support, mental health support, and sabbatical opportunities.
  • We enjoy gathering for meals, cultural initiatives, and events like local Summer Sessions and year‑end celebrations. There are also healthy snacks, drinks, and a weekly catered lunch.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Site Reliability Engineer - Datadog, Kafka (d/f/m)
Site Reliability Engineer - Datadog, Kafka (d/f/m)

Personio • München

Hybrid
EUR 70.000 - 90.000
28 days of paid vacation
Generous family leave and mental health support
Impact Day
+1
Senior Site Reliability Engineer (f/m/d)
Senior Site Reliability Engineer (f/m/d)

Personio • Berlin

Hybrid
EUR 90.000 - 130.000
Competitive reward package including 0
28 days of paid vacation +1 day after
Impact Day
+1
Staff Site Reliability Engineer (d/f/m)
Staff Site Reliability Engineer (d/f/m)

Personio • München

Hybrid
EUR 70.000 - 90.000
Competitive reward package
28 days of paid vacation
Generous family leave and mental health support
+1
Staff Site Reliability Engineer (d/f/m)
Staff Site Reliability Engineer (d/f/m)

Personio • Berlin

Hybrid
EUR 100.000 - 160.000
Competitive reward package
28 days of paid vacation
Impact Day
+5
Staff Site Reliability Engineer (d/f/m)
Staff Site Reliability Engineer (d/f/m)

Devops Academy • München

Hybrid
EUR 70.000 - 90.000
Competitive reward package
28 days of paid vacation
Fully paid Impact Day
+2
Platform Engineer - Infra & Cloud Delivery (f/m/d) , L4
Platform Engineer - Infra & Cloud Delivery (f/m/d) , L4

Personio • Berlin

Hybrid
EUR 90.000 - 130.000
Pre-IPO equity
Office-based with 2 days in office
20 Flex Days per year
+1
Platform Engineer - Infra & Cloud Delivery (f/m/d) , L4
Platform Engineer - Infra & Cloud Delivery (f/m/d) , L4

Personio GmbH • Berlin

Hybrid
EUR 90.000 - 130.000
2 days in office per week
Hybrid work in Berlin/Dublin
Competitive salary
+2
Senior Platform Engineer - Infra & Cloud Delivery (f/m/d) , L5
Senior Platform Engineer - Infra & Cloud Delivery (f/m/d) , L5

Personio • München

Hybrid
EUR 100.000 - 150.000
Testing Platform Engineer - Developer Experience (d/f/m)
Testing Platform Engineer - Developer Experience (d/f/m)

Devops Academy • München

Hybrid
EUR 70.000 - 90.000
28 days of paid vacation
Fully paid Impact Day
Mental health support
+1
Platform Engineer - Developer experience L4 (f/m/d)
Platform Engineer - Developer experience L4 (f/m/d)

Personio • München

Hybrid
EUR 90.000 - 130.000
Equity
28 days vacation
Flex days
+2