Site Reliability Engineer (m/f/d)

Allianz Partners

München

Vor Ort

EUR 90.000 - 120.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Allianz Partners is seeking a Site Reliability Engineer within Advanced Analytics to own the reliability of the central engineering platform. You will define SLOs/SLIs, drive incident response, and automate toil away through Terraform and IaC tooling.

You will work with Platform Engineers, Security Engineers, and incident-response leads to ensure reliability across AI services, Java APIs, and frontend apps.

Qualifikationen

  • 5+ years professional experience in site reliability engineering, DevOps, or platform engineering roles.
  • Strong Kubernetes experience: cluster operations, networking, storage, autoscaling, and troubleshooting.
  • Terraform experience; familiarity with IaC tooling like Bicep or ARM templates is a plus.
  • Production experience with Azure cloud services: AKS, ACR, Key Vault, Azure Monitor, Application Insights.
  • Strong observability experience: Prometheus, Grafana, centralized logging, alerting, and distributed tracing instrumentation.
  • Working knowledge of SLO/SLI methodology: error-budget principles and capacity planning.
  • Structured incident management experience: on-call ownership and blameless post-incident reviews.
  • Scripting and automation proficiency in Python or Bash.
  • Strong CI/CD experience: GitHub Actions and ArgoCD or equivalent GitOps tooling.

Aufgaben

  • Define, instrument, and maintain SLOs and SLIs; own error budget tracking and reliability reports.
  • Lead on-call rotation and incident response for infrastructure issues; chair blameless post-incident reviews.
  • Operate Kubernetes infrastructure (AKS): cluster lifecycle, networking, quotas, autoscaling, multi-tenancy across namespaces.
  • Develop IaC (Terraform) to provision Azure resources with auditability and rollback capability.
  • Build and maintain observability stack: Prometheus, Grafana, Azure Monitor, Application Insights; manage alerts and dashboards.
  • Perform capacity planning and cost-aware resource management; tune autoscalers and right-size node pools.
  • Automate repetitive tasks to reduce toil; track toil reduction over time.
  • Maintain platform reliability procedures: backups, DR runbooks, upgrade strategies, change freezes.
  • Contribute to CI/CD and GitOps tooling (GitHub Actions, ArgoCD) with reliability and safe deployment in mind.
  • Collaborate with incident-response and security teams on SLAs and remediation.

Kenntnisse

Kubernetes
Terraform
Azure
Observability
SRE/DevOps
Python
GitHub Actions
ArgoCD

Tools

AKS
Prometheus
Grafana
Azure Monitor
Application Insights

Jobbeschreibung

Key Responsibilities

As a Site Reliability Engineer within Advanced Analytics (DA3) in the Chief Data & AI Office at Allianz Partners, you will join the platform engineering team to own the reliability and operational health of the central engineering platform.

You will define and maintain service level objectives, drive incident response at the infrastructure layer, and systematically eliminate operational toil through automation.

You will work closely with Platform Engineers, Security Engineers, and incident-response leads to ensure the platform meets its reliability commitments across production workloads spanning AI services, Java APIs, and frontend applications.

Through this role, you will have the main following responsibilities:

  • Define, instrument, and maintain SLOs and SLIs for platform components; own error budget tracking and produce regular reliability reports for senior leadership.
  • Serve on the on-call rotation as the infrastructure escalation tier; lead incident response for cluster-level, network-level, and storage failures; chair blameless post-incident reviews.
  • Implement and operate Kubernetes infrastructure (AKS): cluster lifecycle management, networking, resource quotas, autoscaling configuration, and multi-tenancy patterns across product team namespaces.
  • Develop Infrastructure as Code (Terraform) to provision and manage Azure resources with consistency, auditability, and repeatable rollback capability.
  • Build and maintain observability infrastructure: Prometheus, Grafana, Azure Monitor, and Application Insights; own alerting rules, dashboards, and distributed tracing coverage across platform components.
  • Perform capacity planning and cost-aware resource management: right-size node pools, tune vertical and horizontal pod autoscalers, and identify resource waste across namespaces.
  • Identify and eliminate toil: automate repetitive operational tasks through scripting and tooling; measure and track toil reduction over time.
  • Maintain platform reliability procedures: rolling upgrades, backup and recovery testing, disaster recovery runbooks, and change freeze coordination.
  • Contribute to CI/CD pipelines and GitOps tooling (GitHub Actions, ArgoCD) from a reliability and deployment safety perspective; work with platform engineering on release gates and rollback mechanisms.
  • Collaborate with incident-response leads on incident SLA targets and operational procedures; work with Security Engineers on infrastructure hardening and vulnerability remediation.
What You Bring
  • 5+ years professional experience in site reliability engineering, DevOps, or platform engineering roles.
  • Strong Kubernetes experience: cluster operations, networking (Ingress, network policies), storage, autoscaling, and hands‑on troubleshooting across production environments.
  • Solid Infrastructure as Code experience with Terraform; familiarity with Bicep or ARM templates is a plus.
  • Production experience with Azure cloud services: AKS, ACR, Key Vault, Azure Monitor, Application Insights, Virtual Networks, and Private Endpoints.
  • Strong observability experience: Prometheus, Grafana, centralized logging, alerting configuration, and distributed tracing instrumentation.
  • Working knowledge of SLO/SLI methodology: error‑budget principles, reliability target setting, and capacity planning.
  • Structured incident management experience: on‑call ownership, blameless post‑incident review, and runbook authorship.
  • Scripting and automation proficiency in Python or bash for toil elimination and operational tooling.
  • Strong CI/CD experience: GitHub Actions and ArgoCD or equivalent GitOps tooling.
Ways of Working
  • Comfortable in agile, iterative delivery environments with personal ownership and accountability for platform reliability.
  • Clear communicator across global, cross‑functional stakeholders; able to translate technical reliability metrics into business impact for non‑technical audiences.
  • Proactive learner with pragmatic adoption of AI‑assisted developer tools (e.g., GitHub Copilot, Claude Code) to improve automation coverage and delivery velocity.
Nice to Have
  • Kubernetes certifications: CKA or CKAD.
  • Experience supporting AI or ML infrastructure workloads: GPU scheduling, model serving platforms, or inference pipeline operations.
  • Exposure to chaos engineering practices and fault injection testing.
  • FinOps experience: reserved capacity planning, resource right‑sizing programs, and cost attribution per team or workload.
  • Service mesh experience (Istio, Linkerd) for traffic management and reliability patterns.
  • Experience in regulated industries (insurance, finance, healthcare) where auditability, change traceability, and secure‑by‑default operations are standard practice.
What We Offer

Our employees play an integral part in our success as a business. We appreciate that each of our employees are unique and have unique needs, ambitions and we enjoy being a part of their journey. We are there to empower and encourage you with your personal and professional development ensuring that you take control by offering a large variety of courses and targeted development programs.

All that in a global environment where international mobility and career progression are encouraged. Caring for your health and wellbeing is key priority for us. This is why we build Work Well programs to providing you with peace of mind and give the flexibility in planning and arranging for a better work‑life balance.

90377 | Data & AI | Professional | Allianz Partners | Full‑Time | Permanent

We therefore welcome applications regardless of ethnicity or cultural background, age, gender, nationality, religion, social class, disability or sexual orientation, or any other characteristics protected under applicable local laws and regulations.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Platform Engineer (Cloud Infrastructure, AI Platform) (m/f/d)
Platform Engineer (Cloud Infrastructure, AI Platform) (m/f/d)

Allianz Partners • München

Vor Ort
EUR 65.000 - 85.000
Work Well programs
Career development courses
International mobility opportunities
Senior Backend Engineer (m/f/d)
Senior Backend Engineer (m/f/d)

Allianz Partners • München

Vor Ort
EUR 70.000 - 90.000
Training and development programs
Work-life balance initiatives
Engineering Manager (Infrastructure)
Engineering Manager (Infrastructure)

Allianz Partners • München

Vor Ort
EUR 130.000 - 170.000
AI Platform Engineer (m/f/d)
AI Platform Engineer (m/f/d)

Tamarind Intelligence • München

Hybrid
EUR 90.000 - 130.000
Hybrid work model (incl. up to 25 days
Company bonus scheme
Pension
+2
AI Engineer (m/f/d)
AI Engineer (m/f/d)

Allianz • Unterföhring

Hybrid
EUR 70.000 - 100.000
Hybrid work model
Company bonus scheme
Pension
+2
Engineering Manager (Backend)
Engineering Manager (Backend)

Allianz Partners • München

Vor Ort
EUR 120.000 - 180.000
AI Engineer (m/f/d)
AI Engineer (m/f/d)

Allianz Technology • Unterföhring

Hybrid
EUR 70.000 - 110.000
Company bonus scheme
Pension
Employee shares program
+1
Head of Core Engineering
Head of Core Engineering

Allianz Partners • München

Vor Ort
EUR 150.000 - 210.000
Engineering Manager AI
Engineering Manager AI

Allianz Partners • München

Vor Ort
EUR 110.000 - 170.000
Intern (m/f/d) - Allianz Services - Data and AI Transformation
Intern (m/f/d) - Allianz Services - Data and AI Transformation

Allianz Services • Unterföhring

Vor Ort
EUR 17.000 - 20.000
Modern Munich office
Flexible hours
LinkedIn Learning
+3