Senior Site Reliability Engineer (m/w/d)

Impower

Germany (OH)

On-site

USD 81,603 - 104,919

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible working hours
Real ownership in projects
Growth opportunities in a modern environment

Job summary

Impower, based in Munich, is seeking a Senior Site Reliability Engineer to enhance our AI-driven ERP platform's reliability and operational foundations. You will work with technologies like Kubernetes and AWS to ensure a scalable and resilient system.

Candidates should have 5+ years of cloud engineering experience, deep Kubernetes knowledge, and strong AWS skills. This role offers a hybrid work setup with opportunities for growth in a supportive team environment.

Qualifications

  • 5+ years building and operating production systems in cloud environments.
  • Deep, hands-on production Kubernetes experience.
  • Strong working knowledge of EKS, RDS, ALB, IAM, VPC, S3.
  • Solid Terraform experience with disciplined IaC practices.
  • Hands-on experience with ArgoCD, Helm, or equivalent tools.
  • IAM best practices knowledge and compliance frameworks familiarity.
  • Experience building dashboards, defining SLOs, and incident management.
  • Comfortable scripting and building tooling in Python, Go, Bash.
  • Excellent written and verbal English (C1+).

Responsibilities

  • Own platform reliability end-to-end.
  • Drive CI/CD excellence with GitLab + Terraform + ArgoCD.
  • Design and operate scalable AWS infrastructure.
  • Improve observability with Sentry, Grafana, Prometheus.
  • Own security posture across infrastructure layers.
  • Strengthen on-call practices and post-incident learning.
  • Provide tooling and guidance to product engineers.
  • Support broader platform surface as AI capabilities expand.

Skills

Kubernetes depth
AWS expertise
Infrastructure as Code
GitOps
Security expertise
Observability instincts
Automation fluency
Communication

Education

5+ years engineering experience in cloud environments

Tools

Terraform
GitLab
ArgoCD
Helm

Job description

Introduction

Munich | Hybrid | Product based company

At Impower, we are shaping the future ofproperty management — simple, fast, and digital. More than 100,000 people manage over 12 million apartments using processes that are still manual, complex, and time-consuming. We help property managers modernize these workflows step by step with reliable software, scalable systems, and real practical value.

As our platform continues to grow, reliability, scalability, security, and developer enablement become increasingly critical. We are building an AI-driven ERP platform that combines modern cloud infrastructure, workflow automation, and intelligent services to support complex operational processes at scale.

As aSenior Site Reliability Engineer, you will take ownership of the reliability and operational foundations of our platform. You will work across Kubernetes, AWS infrastructure, CI/CD, observability, and security to ensure that our systems remain scalable, resilient, and secure as we grow.

Your mission
  • Own platform reliability end-to-end:Co-own our Kubernetes-based platform on AWS alongside our current Senior SRE - ingress, autoscaling, service mesh, config and secrets — and make sure it scales as we grow.

  • Drive CI/CD excellence:Evolve our GitLab + Terraform + ArgoCD/Helm pipelines to deliver our Java/Spring Boot and React applications faster, safer, and with more self-service capability for product teams.

  • Manage cloud infrastructure:Design and operate scalable AWS infrastructure (EKS, RDS, ALB, IAM, VPC, S3) using Infrastructure as Code, with strong IaC discipline and clear change management.

  • Strengthen observability:Improve our Sentry, Grafana, Prometheus, and Loki setup so teams can define SLOs, debug fast, and operate their services with confidence.

  • Lead on security:Own our security posture across infrastructure and application layers — IAM, secrets management, network segmentation, container and dependency scanning, vulnerability management, supply chain security, and audit readiness. Embed security as a design constraint, not a bolted-on review step.

  • Improve incident response:Strengthen our on-call practices, runbooks, and post-incident learning. We treat reliability as a product feature.

  • Enable product teams:Provide tooling, guidance, and self-service capabilities that help product engineers adopt better operational and deployment practices — make the good path the easy path.

  • Support the broader platform surface:Temporal workflows, PostgreSQL operations, S3, our Estuary CDC pipeline, and AI service infrastructure on GCP/Azure as we expand our AI capabilities

Your profile
  • Engineering experience:5+ years building and operating production systems in cloud environments, including real ownership of non-trivial systems at scale.

  • Kubernetes depth:Deep, hands-on production Kubernetes experience — beyond kubectl apply, including operators, networking, autoscaling, and debugging.

  • AWS expertise:Strong working knowledge of EKS, RDS, ALB, IAM, VPC, S3, and the operational realities of running services on AWS.

  • Infrastructure as Code:Solid Terraform experience with disciplined IaC practices.

  • GitOps:Hands-on experience with ArgoCD, Helm, or equivalent declarative deployment tooling.

  • Security expertise applied to cloud-native environments:IAM best practices, secrets management, secure network architecture, container and dependency vulnerability scanning, secure SDLC principles, and familiarity with compliance frameworks (e.g. ISO 27001, SOC 2, or comparable). You proactively identify risks and contribute to incident response and audit readiness.

  • Observability instincts:You've built dashboards, defined SLOs, run real incidents, and used the resulting learning to improve systems.

  • Automation fluency:Comfortable scripting and building tooling in Python, Go, Bash, or similar.

  • Communication:Excellent written and verbal English (C1+). You document decisions, write runbooks people actually use, and explain tradeoffs clearly.

Nice to have
  • Experience with Temporal or other workflow orchestration systems

  • Exposure to CDC pipelines (Estuary, Debezium, or similar)

  • Multi-cloud experience (AWS primary, GCP/Azure for AI services)

  • Background running platforms for Spring/Java services at scale

  • Experience in regulated environments (financial services, real estate, healthcare)

  • Familiarity with AI/LLM infrastructure patterns or agentic engineering workflows

  • Prior experience as an early platform hire — building foundations without over-engineering

How you work
  • You think in systems and prefer building self-service capabilities over becoming a ticket queue

  • You're pragmatic about quality — protecting long-term adaptability without gold-plating

  • You communicate openly about tradeoffs, mistakes, and unknowns.

  • You see security and compliance as enablers of speed, not obstacles.

  • You're comfortable with autonomy and high ownership in a small, focused team.

  • You're curious about how AI is changing platform engineering — and you want to help figure it out.

Why us?

Freedom and flexibility:Hybrid setup (3 days / week in the office) from Munich, with flexible hours and real ownership.

Meaningful impact:Build technology that helps thousands of people simplify property management every day.

Modern environment:Work with a cutting-edge tech stack (React, TypeScript, Java, AWS, Kubernetes) and up-to-date tools.

Growth opportunities:Deep technical scope across SRE, cloud infrastructure, security, and automation — shaping the reliability and security foundations of the platform as we scale.

Supportive culture:Join a diverse, agile team that values autonomy, trust, and collaboration, with strong guidance during onboarding and beyond.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Juniper Square • United States

Remote
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+4
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

FACT-Finder • Germany (OH)

Hybrid
USD 103,000 - 139,000
Hybrid work model
Open feedback culture
Reliability discipline
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

Juniper Square • United States

On-site
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+3
Staff Core Platform Engineer
Staff Core Platform Engineer

United States Digital Space LLC • United States

Remote
USD 180,000 - 240,000
Equity
Remote-first
Open source stipend
Senior Software Engineer, Platform (Security Focused) (m/f/d)
Senior Software Engineer, Platform (Security Focused) (m/f/d)

casavi GmbH • Germany (OH)

Hybrid
USD 103,000 - 144,000
Flexible hybrid work
Workation 40 days
Company events
+5
Senior Site Reliability Engineer
Senior Site Reliability Engineer

hardrockdigital • United States

Hybrid
USD 120,000 - 160,000
Competitive pay and benefits
Flexible vacation allowance
Startup culture with global brand support
+1
Director of Platform Infrastructure (VP of Infrastructure / Head of Site Reliability Engineerin[...]
Director of Platform Infrastructure (VP of Infrastructure / Head of Site Reliability Engineerin[...]

Naviga Recruiting & Executive Search • Tampa (FL)

Hybrid
USD 200,000 - 260,000
Hybrid work
Global travel
Staff Site Reliability Engineer (x/f/m)
Staff Site Reliability Engineer (x/f/m)

United States Digital Space LLC • United States

Hybrid
USD 127,000 - 161,000
Deutschlandticket (Germany-wide public
28 vacation days
Work from abroad up to 10 days/year
+9
Senior DevOps Engineer - LATAM (Remote)
Senior DevOps Engineer - LATAM (Remote)

Socket.dev • United States

Remote
USD 140,000 - 210,000
Senior DevOps Engineer – LATAM
Senior DevOps Engineer – LATAM

Luxury Presence • United States

Remote
USD 140,000 - 210,000