Senior Site Reliability Engineer (SRE / Backend) f/m/d

DUDE CHEM

Berlin

Hybrid

EUR 110.000 - 170.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Real ownership
nilo app access
Learning budget
Work abroad up to 90 days
Hybrid work 2/3
Urban Sports Club
Equity options
Team events
Dog-friendly office

Zusammenfassung

nilo is seeking a Senior Site Reliability Engineer to own the AWS-based infrastructure and drive reliability for web and mobile applications. You will manage infrastructure as code with Terraform, design resilient, event-driven services, and lead incident response.

This hands-on role also touches backend work and security, in a hybrid office setting in Berlin. You will collaborate with product and security teams to implement scalable operator dashboards, SLOs, and cost controls, while mentoring

Qualifikationen

  • 5+ years in SRE/DevOps with production ownership
  • Deep AWS experience with serverless and containers
  • Strong Terraform skills and module design
  • Experience with event-driven architectures and failure modes
  • PostgreSQL tuning, indexing, and migrations
  • Observability with Datadog or similar tools
  • Production backend coding in Python/Node.js/Go
  • On-call and incident postmortems experience
  • Security fundamentals: IAM, secrets, network isolation

Aufgaben

  • Own AWS infrastructure end to end (Lambda, ECS Fargate, SQS, SNS, EventBridge, SES, Cognito, DynamoDB, RDS Postgres, DMS)
  • Manage infrastructure as code in Terraform with modular design
  • Build and maintain CI/CD pipelines with safe rollout and rollback
  • Design event-driven services for resilience and idempotency
  • Own monitoring dashboards and alerting (Datadog, SLOs)
  • Lead incident response and blameless postmortems
  • Harden security posture and GDPR/compliance readiness
  • Monitor and optimize AWS spend without compromising reliability
  • Contribute to backend development and data pipelines
  • Mentor engineers on operational excellence

Kenntnisse

SRE ownership
Incident response
On-call reliability
Excellent communication

Tools

Terraform
Datadog
Grafana
PostgreSQL
AWS

Jobbeschreibung

About the Role

We're looking for a Senior Site Reliability Engineer to take ownership of the infrastructure behind our platform. We run web applications and two mobile apps entirely on AWS, built around serverless and event-driven services and managed with Terraform.

This is a hands‑on role with real ownership — you'll set the standards for how we run production rather than inherit someone else's. The split is roughly 70% infrastructure and reliability work, 30% backend development. Most of your time goes to the platform, but you'll be comfortable dropping into the application code to debug a slow query, fix a Lambda, or ship an endpoint alongside the product team.

Because we handle sensitive mental health data, reliability and security aren't abstract goals here. When someone books a session in a difficult moment, our platform needs to work.

What you'll do
  • Own our AWS infrastructure end to end — Lambda, ECS Fargate, SQS, SNS, EventBridge, SES, Cognito, DynamoDB, RDS Postgres, and DMS

  • Manage everything as code in Terraform, with well-designed modules, clean state management, and a solid review workflow

  • Build and maintain CI/CD pipelines with safe rollout and rollback across web, mobile backends, and infrastructure

  • Design our event-driven services for resilience: retries, dead‑letter queues, idempotency, graceful degradation

  • Own our Datadog and Sentry setup — define SLOs, build dashboards, and keep alerting actionable instead of noisy

  • Lead incident response and run blameless postmortems that actually change how we build

  • Harden our security posture: IAM, secrets management, network boundaries, Cognito auth flows, and vulnerability remediation

  • Protect sensitive health data and support our GDPR and compliance requirements

  • Monitor and optimize AWS spend without compromising reliability

  • Contribute to backend development — APIs, event consumers, data pipelines, and Postgres and DynamoDB performance

  • Participate in architectural discussions and mentor engineers on operational excellence

What we're looking for
Must-have
  • 5+ years in SRE, DevOps, platform, or backend engineering, with real production ownership

  • Deep AWS experience across serverless and containers — Lambda, ECS Fargate, and debugging both under pressure

  • Strong Terraform skills, including module design and managing state across multiple environments

  • Hands‑on experience with event‑driven architecture (SQS, SNS, EventBridge) and a healthy respect for its failure modes

  • Solid PostgreSQL: query tuning, indexing, connection management, and zero‑downtime migrations

  • Production experience with Datadog or a comparable observability platform (Grafana, New Relic, Honeycomb)

  • Comfortable writing production backend code in [Python / Node.js / Go]

  • Genuine on‑call and incident response experience — you've led an incident and written the postmortem

  • Strong security fundamentals: IAM, least privilege, secrets, network isolation, common web vulnerabilities

  • Pragmatic about complexity — you reach for the simplest thing that meets the reliability bar

  • Effective communicator who can explain a technical trade‑off without jargon

Nice‑to‑have
  • AWS DMS or other data migration and replication tooling

  • Compliance experience (GDPR, SOC 2, ISO 27001)

  • Experience in a B2B SaaS environment or healthcare‑related product

What you get
  • Real ownership: a small team, short feedback loops, and no layers of approval between you and production

  • Work that matters: the reliability you build directly affects people reaching for mental health support

  • Free access to the nilo app (incl. family support)

  • A dedicated learning budget for your personal and professional development

  • Work abroad for up to 90 days per year (within the EU)

  • Hybrid working model: 2 days/week from the office, 3 days from home

  • Urban Sports Club membership at a discounted price

  • Equity options: you benefit from any increase in nilo's valuation that you've helped to create

  • Regular team and company events

  • Bring your dog to work: we have 4 office dogs

nilo embraces diversity. We strive to create an inclusive workplace where everyone feels welcome, psychologically and physically safe. All applicants will be considered for employment without regard to different ethnic/racial origins, age groups, religions/ideologies, sexual orientations, gender identities, abilities, socio-economic statuses, educational backgrounds, family arrangements and/or any other characteristic.

We would like to encourage you to apply even if the technical requirements cannot be met 100%.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer (SRE / Backend) f/m/d
Senior Site Reliability Engineer (SRE / Backend) f/m/d

nilo • Berlin

Hybrid
EUR 110.000 - 150.000
Real ownership in a small team
Mental health platform access for you/
Free nilo app access (incl. family)
+7
Senior Site Reliability Engineer (SRE / Backend) f/m/d
Senior Site Reliability Engineer (SRE / Backend) f/m/d

Atlantic Labs • Berlin

Hybrid
EUR 90.000 - 120.000
Real ownership
Nilo app access
Learning budget
+6
Senior Site Reliability Engineer (m/f/d)
Senior Site Reliability Engineer (m/f/d)

TOPdesk • Kaiserslautern

Vor Ort
EUR 90.000 - 150.000
30 days annual vacation
Remote-friendly options
Health and wellness programs
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

B Capital • Deutschland

Remote
EUR 46.000 - 105.000
Work from anywhere
Flexible paid time off
Mental health support services
+3
Senior Site Reliability Engineer (f/m/d)
Senior Site Reliability Engineer (f/m/d)

Personio • Berlin

Hybrid
EUR 90.000 - 130.000
Competitive reward package including 0
28 days of paid vacation +1 day after
Impact Day
+1
Senior Site Reliability Engineer (m/f/d) at TOPdesk
Senior Site Reliability Engineer (m/f/d) at TOPdesk

TOPdesk • Kaiserslautern

Vor Ort
EUR 90.000 - 120.000
Possibility to work remote
Flexible working hours
Lead Backend Engineer (m/f/d)
Lead Backend Engineer (m/f/d)

Peter Park • München

Hybrid
EUR 60.000 - 80.000
Competitive salary
Flexible working hours
Health & wellness benefits
+2
Head of Engineering - Unified Data Platform
Head of Engineering - Unified Data Platform

Nelly Solutions • Berlin

Hybrid
EUR 140.000 - 210.000
BVG ticket or Swapfiets membership
Urban Sports membership
Top-of-the-line equipment
+7
Staff Site Reliability Engineer (d/f/m)
Staff Site Reliability Engineer (d/f/m)

Devops Academy • München

Hybrid
EUR 70.000 - 90.000
Competitive reward package
28 days of paid vacation
Fully paid Impact Day
+2
Staff Site Reliability Engineer (d/f/m)
Staff Site Reliability Engineer (d/f/m)

Personio • Berlin

Hybrid
EUR 100.000 - 160.000
Competitive reward package
28 days of paid vacation
Impact Day
+5