Senior Site Reliability Engineer / Platform Engineer

SimScale

München

Vor Ort

EUR 70.000 - 90.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Mobile Working
Competitive Health Benefits
Discounted Gym Membership
Tech Talks and Social Events
Flexible Working Hours
Learning and Development
Child Care Contributions
Retirement Plan

Zusammenfassung

SimScale is seeking a Senior SRE / Platform Engineer in Munich to enhance the cloud infrastructure of its renowned simulation platform. This role involves building standards and tooling for AWS, driving observability with OpenTelemetry, and managing a scalable multi-region architecture.

The ideal candidate has a strong systems background, software development skills in Python, Go, Rust, or Java, and a commitment to reliability and security. Flexible working hours and competitive benefits are offered to support a healthy work-life balance.

Qualifikationen

  • Strong understanding of Linux internals and distributed systems.
  • Experience in software development and SRE.
  • Knowledge of security compliance standards like SOC 2.
  • Ability to debug complex production issues.
  • Hands-on experience with AWS and Kubernetes.
  • Prior technical leadership experience.
  • Experience in observability and reliability practices.

Aufgaben

  • Develop standards and tooling for AWS workloads.
  • Support a team of 50+ engineers in infrastructure.
  • Evolve the Kubernetes platform and evaluate new technologies.
  • Drive adoption of OpenTelemetry for observability.
  • Shape multi-cloud architecture for data residency.
  • Manage cloud cost and efficiency.

Kenntnisse

Linux internals
Software development (Python, Go, Rust, Java)
Security and compliance
Cloud and infrastructure (AWS, GCP)
Observability (OpenTelemetry, Prometheus)
5+ years in SRE/Platform Engineering
Clear communication

Tools

Terraform
Kubernetes
ArgoCD

Jobbeschreibung

We are looking for a Senior SRE / Platform Engineer (m/f/d) to own and improve the cloud infrastructure behind SimScale’s browser-based simulation platform. The role spans AWS and EKS, observability, disaster recovery, security and compliance controls, multi-region architecture, elastic GPU/HPC capacity, and internal developer tooling.

Responsibilities
  • SimScale’s engineering teams run workloads directly on AWS; you will build the standards, guardrails, and self-service tooling that let them do so safely, raising reliability and security without slowing engineering velocity.
  • You will join a small, tightly knit infrastructure team supporting 50+ engineers across the company. This is a hands‑on senior individual contributor role; people management is not required, but there is a genuine path toward tech‑lead ownership as the team grows.
  • Evolve our Kubernetes platform: Evaluate and adopt technologies such as Kubernetes Gateway API and service mesh patterns, and coordinate platform evolution across 10+ engineering teams.
  • Take observability to the next level: Drive organization-wide adoption of OpenTelemetry for distributed tracing and metrics, and help teams define meaningful SLOs.
  • Shape multi-region architecture and data residency: Support our move from an EU-centered footprint toward a global, multi-cloud architecture that satisfies disaster‑recovery and data‑residency requirements.
  • Own cloud cost and efficiency at scale: Keep petabyte-scale infrastructure cost‑efficient, secure, and well‑instrumented.
  • Improve tooling: Build self‑service AWS account provisioning, guardrails and AI‑assisted automations that help engineering teams manage infrastructure safely and efficiently at scale.
Qualifications
  • Strong systems foundation: You understand Linux internals and distributed systems well enough to debug complex production behavior.
  • Software development experience: Your background is rooted in software development, and you moved into SRE from there. You write production-quality software in at least one of Python, Go, Rust, or Java.
  • Security and compliance awareness: You understand how infrastructure decisions affect access control, auditability, disaster recovery, logging, and standards such as SOC 2.
  • Production debugging depth: You can investigate complex failures, communicate clearly during incidents, and turn findings into durable improvements.
  • Hands‑on cloud and infrastructure experience: AWS (or GCP), declarative infrastructure (Terraform), gitops‑workflow (ArgoCD) and container orchestration (Kubernetes).
  • 5+ years of professional experience in SRE, platform, or infrastructure engineering.
  • Clear communication: You can explain trade‑offs to engineering teams and help others adopt better platform practices without unnecessary friction.
  • Observability and reliability experience: You have worked with OpenTelemetry, Prometheus, distributed tracing, monitoring, and meaningful SLOs/SLIs.
  • An open source portfolio or contributions.
  • Prior technical leadership experience, especially in infrastructure, reliability, or platform engineering.
Benefits
  • Mobile Working: Modern technology coupled with the widespread use of digital communication enables remote working. We embrace mobile working as it offers many possibilities for our employees.
  • Competitive Health Benefits: Whilst working in a competitive and ambitious environment, it is just as important to take care of your health. At SimScale, we’ve got you covered.
  • Discounted Gym Membership: We offer a gym membership with multiple locations to make it easier for our employees to invest in their health, well‑being, and workout regularly.
  • Tech Talks and Social Events: Whether it’s through tech talks, social events or other organized activities, we are not just working together, we like sharing, and enjoying each other’s company.
  • Flexible Working Hours: It doesn’t matter if you are an early bird or a night owl, at SimScale you have the freedom to plan your workday.
  • Learning and Development: With the right training in place, we offer our employees the opportunity to further develop their competencies and skill-sets, to ultimately become more successful and satisfied.
  • Child Care Contributions: In addition to flexible working hours, and home office possibilities to support family commitments, we also provide contributions for our mini SimScaler’s nursery or kindergarten childcare.
  • Retirement Plan: We are future-oriented, and take care of our employees! We provide employees with a retirement plan that will contribute towards a comfortable and secure future.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Platform Engineer (AWS)
Senior Platform Engineer (AWS)

GotPhoto • Berlin

Hybrid
EUR 70.000 - 90.000
Unlimited paid holiday
Education budget
Remote work options
+1
Senior Site Reliability Engineer (m/w/d)
Senior Site Reliability Engineer (m/w/d)

Impower • München

Hybrid
EUR 70.000 - 90.000
Flexible hours
Ownership in projects
Diverse team culture
Customer Support Agent EMEA (m/f/d)
Customer Support Agent EMEA (m/f/d)

SimScale • Köln

Remote
Customer Success Manager (m/f/d)
Customer Success Manager (m/f/d)

SimScale GmbH • München

Remote
Vertraulich
Autonomy to build strategy
Opportunity to work with cutting-edge technology
Focus on diversity, equity, and inclusion
Senior Site Reliability Engineer (m/f/d)
Senior Site Reliability Engineer (m/f/d)

TOPdesk • Kaiserslautern

Vor Ort
EUR 90.000 - 150.000
30 days annual vacation
Remote-friendly options
Health and wellness programs
+2
Senior Site Reliability Engineer (SRE / Backend) f/m/d
Senior Site Reliability Engineer (SRE / Backend) f/m/d

nilo • Berlin

Hybrid
EUR 110.000 - 150.000
Real ownership in a small team
Mental health platform access for you/
Free nilo app access (incl. family)
+7
(Senior) Cloud Site Reliability Engineer (Scalability) (m/f/x)
(Senior) Cloud Site Reliability Engineer (Scalability) (m/f/x)

Scalable Capital • Berlin

Hybrid
EUR 70.000 - 90.000
Flexible vacation policy
Monthly contribution for the Deutschland Jobticket
Discounted sports activities
Platform Engineer
Platform Engineer

Aether Biomedical • Deutschland

Hybrid
EUR 90.000 - 120.000
Vacation days up to 26
Illness days: 10 per year
Health and life insurance
+6
Senior Site Reliability Engineer (m/f/d) at TOPdesk
Senior Site Reliability Engineer (m/f/d) at TOPdesk

TOPdesk • Kaiserslautern

Vor Ort
EUR 90.000 - 120.000
Possibility to work remote
Flexible working hours
Senior Platform Engineer | Climate Tech startup
Senior Platform Engineer | Climate Tech startup

Carbmee GmbH • Berlin

Hybrid
EUR 90.000 - 120.000
Stock options
Hybrid office in Berlin/Munich
Health & wellbeing programs
+1