Reliability & Observability Engineer (m/f/d)

Personio

München

Hybrid

EUR 90.000 - 140.000

Vollzeit

vor 48 Stunden
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Eine Bewerbung wie gemacht für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Hybrid work culture (Munich office)
Flexible hours
26+4 vacation days
30 days of workation in EU
Autonomy and flat hierarchies
Udemy access
EGYM Wellpass

Zusammenfassung

Personio is seeking an experienced software platform engineer to own reliability for our regulatory crawler stack. You will write production Python, build agentic incident responses, and improve observability and automation across the pipeline in a hybrid Munich-based role.

We value hands-on experience with LLMs, GCP, containers, and CI/CD, and you will help reduce outages, improve data freshness, and drive cost-efficient infrastructure while collaborating in a small, mission-driven team.

Qualifikationen

  • 5+ years in software engineering, SRE or platform roles with production Python experience.
  • Proven ability to turn alerts into actionable work.
  • Experience with LLMs or agents in production.
  • Strong knowledge of GCP, containers and CI/CD.
  • Understanding of distributed-system failure modes.

Aufgaben

  • Reduce noise by distinguishing transient from real failures and build reliable alerts.
  • Build agentic incident response with Sentry issues, logs, metrics and deploys.
  • Monitor data quality: freshness and completeness across sources.
  • Make the pipeline self-healing with retries, idempotent jobs and dead-letter handling.
  • Ensure safe agent actions with scoped permissions and human approvals.
  • Audit infrastructure for cost savings and efficiency opportunities.
  • Catch memory, timeout and cost issues across Cloud Run.
  • Add structured logging, metrics, tracing and Sentry grouping.

Kenntnisse

Python
GCP
Containers
CI/CD
LLMs

Ausbildung

5+ years in software engineering

Tools

Docker
Sentry
MongoDB
Cloud Run
Go
React

Jobbeschreibung

Company Culture
About the Job

Our crawlers collect regulatory updates from 80+ sources and push them through an automated pipeline: extraction, embedding, search, classification, consolidation and translation. As we add regions, keeping it healthy has become a job of its own. We want an engineer who makes production tell us what's wrong before customers do, fixes what can be fixed automatically, and builds AI agents that diagnose the rest. When an issue does reach a developer, it should arrive with a root cause and a suggested fix. Not a ticket-driven ops role: you'll write production Python from week one and own platform reliability alongside a core-team developer

What you will do
  • Cut the noise. Separate transient failures (network blips, timeouts, errors that vanish on rerun) from real ones, and build alerting the team trusts.
  • Build agentic incident response: agents that gather Sentry issues, logs, metrics and recent deploys, classify the failure, apply known fixes or open draft PRs, and brief the right
    developer.
  • Monitor the data, not just the infrastructure. A job that succeeds but extracts nothing is still a failure. Track freshness and completeness, e.g. \"every source checked on time\", \"every
    document produced text and embeddings\".
  • Make the pipeline self-healing: retries with backoff, idempotent and resumable jobs, deadletter handling, clear escalation when automation gives up.
  • Keep agents safe: scoped permissions, audit trails, human approval for risky actions, and
    measuring how often they're right.
  • Continuously audit our Infrastructure and Identify opportunities to make it more efficient and save costs.
  • Catch memory, timeout and cost problems across Cloud Run before they become silent
    OOM kills.
  • Add structured logging, metrics, tracing and sensible Sentry grouping, with infrastructure as code and CI/CD.
Our Tech Stack

Python, Docker, GCP (Cloud Run, Cloud Logging, Cloud Monitoring), Sentry, MongoDB, Azure Blob, LLM APIs. Some Go and React.

Your profile
  • 5+ years in software engineering, SRE or platform roles, with real production Python.
  • A track record of turning an ignored alert channel into one people act on.
  • Hands-on experience building with LLMs or agents in production, and judgment about when
    a plain if-statement is better.
  • Solid GCP (or similar), containers and CI/CD.
  • Strong grasp of distributed-system failure modes: retries, idempotency, partial failure,
    backpressure.
  • Pragmatism, and clear communication in a small team.
Nice to have

OpenTelemetry, SLOs, Terraform or Pulumi, data-quality observability, crawling or document processing.

Why us?
  • Hybrid work culture:join us in our Munich office (min. 2 days/week).
  • Flexibleworking hours.
  • 26+4 vacation daysper year (4 fixed “company rest days” over Christmas).
  • 30 days of “workation”per year, within the EU and selected countries.
  • High autonomyand flat hierarchies.
  • EGYM Wellpassfor unlimited access to fitness courses and gyms.
  • Udemyaccess for educational videos.
Closing

We welcome candidates from all backgrounds and encourage diversity in our team. We encourage female and diverse engineers to apply and join our mission-driven culture that values open communication, work-life balance, and a welcoming environment. Convince us with your personality and your skills, and together we will make great things happen!

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Site Reliability Engineer (m/w/d)
Site Reliability Engineer (m/w/d)

PAYBACK • München

Vor Ort
EUR 55.000 - 75.000
Delicious meals in canteen
24/7 access to gym
Flexible working hours
+3
Full-Stack Engineer | Web Crawling (m/f/d)
Full-Stack Engineer | Web Crawling (m/f/d)

Personio • München

Hybrid
EUR 90.000 - 120.000
60% work from home option
Flexible working hours
30 vacation days per year
+3
Senior Site Reliability Engineer (f/m/d)
Senior Site Reliability Engineer (f/m/d)

Personio • Berlin

Vor Ort
EUR 90.000 - 130.000
Competitive reward package including 0
28 days of paid vacation +1 day after
Impact Day
+1
Senior Software Engineer, Reliability (f/m/d)
Senior Software Engineer, Reliability (f/m/d)

Personio • München

Hybrid
EUR 90.000 - 150.000
Staff Site Reliability Engineer (d/f/m)
Staff Site Reliability Engineer (d/f/m)

Personio • Berlin

Vor Ort
EUR 100.000 - 160.000
Competitive reward package
28 days of paid vacation
Impact Day
+5
Staff Site Reliability Engineer (d/f/m)
Staff Site Reliability Engineer (d/f/m)

Devops Academy • München

Vor Ort
EUR 70.000 - 90.000
Competitive reward package
28 days of paid vacation
Fully paid Impact Day
+2
Staff Site Reliability Engineer (d/f/m)
Staff Site Reliability Engineer (d/f/m)

Personio • München

Vor Ort
EUR 70.000 - 90.000
Competitive reward package
28 days of paid vacation
Generous family leave and mental health support
+1
TechOps Engineer
TechOps Engineer

Silversmith Capital Partners • Berlin

Hybrid
EUR 90.000 - 120.000
Annual learning budget
30 days annual leave
Home office setup budget
+7
Staff Software Engineer - Reliability (d/f/m)
Staff Software Engineer - Reliability (d/f/m)

Personio • Deutschland

Hybrid
EUR 110.000 - 150.000
Competitive reward package
28 days paid vacation
Impact Day
+2
Senior DevOps / Platform Engineer (F/M/X)
Senior DevOps / Platform Engineer (F/M/X)

Project J Ltd • Berlin

Vor Ort
EUR 95.000 - 130.000
Paid Time Off - 30 vacation days
Competitive salary
Training & Development budget €1,500/€
+2