Senior Site Reliability Engineer (SRE / Backend) f/m/d

Atlantic Labs

Berlin

Hybrid

EUR 90.000 - 120.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Mach aus dieser Rolle ein Vorstellungsgespräch — ein Lebenslauf und ein Anschreiben, die darauf ausgerichtet sind, was dieser Arbeitgeber sucht.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Real ownership
Nilo app access
Learning budget
EU travel up to 90 days
Hybrid work 2/3
Urban Sports Club
Equity options
Team events
Dog-friendly office

Zusammenfassung

Atlantic Labs is seeking a Senior Site Reliability Engineer to own the infrastructure behind our platform. We run web apps and mobile apps on AWS with serverless and Terraform, balancing reliability with speed.

This hands-on role splits time 70% infra/reliability and 30% backend development. You’ll design resilient event-driven services, lead incident postmortems, and ensure GDPR compliance while optimizing AWS spend.

Qualifikationen

  • 5+ years in SRE/DevOps with production ownership.
  • Deep AWS experience across serverless and containers.
  • Strong Terraform skills, including module design and multi-environment state management.
  • Hands-on experience with event-driven architecture (SQS, SNS, EventBridge).
  • Solid PostgreSQL: tuning, indexing, migrations.
  • Production observability experience (Datadog, Grafana, etc).
  • Proficient backend coding in Python/Node.js/Go.
  • On-call and incident response leadership.
  • Strong security fundamentals: IAM, least privilege, secrets, network isolation.
  • Clear communicator able to explain tradeoffs without jargon.

Aufgaben

  • Own AWS infrastructure end to end (Lambda, ECS Fargate, SQS, SNS, EventBridge, SES, Cognito, DynamoDB, RDS).
  • Manage infrastructure as code with Terraform and modular design.
  • Build and maintain CI/CD pipelines with safe rollout and rollback.
  • Design resilient event-driven services (retries, DLQs, idempotency).
  • Own Datadog and Sentry setup; define SLOs and dashboards.
  • Lead incident response and postmortems to drive changes.
  • Harden security posture: IAM, secrets, network boundaries.
  • Protect GDPR/compliance requirements.
  • Monitor and optimize AWS spend without sacrificing reliability.
  • Contribute to backend development (APIs, data pipelines, Postgres, DynamoDB).
  • Participate in architectural discussions; mentor engineers.

Kenntnisse

SRE ownership
AWS serverless
Terraform
Event-driven architecture
PostgreSQL
Observability
Backend coding
On-call incidents
Security fundamentals
Communication

Tools

Datadog
Grafana
New Relic
Honeycomb

Jobbeschreibung

About the Role

We're looking for a Senior Site Reliability Engineer to take ownership of the infrastructure behind our platform. We run web applications and two mobile apps entirely on AWS, built around serverless and event-driven services and managed with Terraform.

This is a hands-on role with real ownership - you'll set the standards for how we run production rather than inherit someone else's. The split is roughly 70% infrastructure and reliability work, 30% backend development. Most of your time goes to the platform, but you'll be comfortable dropping into the application code to debug a slow query, fix a Lambda, or ship an endpoint alongside the product team.

Because we handle sensitive mental health data, reliability and security aren't abstract goals here. When someone books a session in a difficult moment, our platform needs to work.

What you’ll do
  • Own our AWS infrastructure end to end - Lambda, ECS Fargate, SQS, SNS, EventBridge, SES, Cognito, DynamoDB, RDS Postgres, and DMS
  • Manage everything as code in Terraform, with well-designed modules, clean state management, and a solid review workflow
  • Build and maintain CI/CD pipelines with safe rollout and rollback across web, mobile backends, and infrastructure
  • Design our event-driven services for resilience: retries, dead-letter queues, idempotency, graceful degradation
  • Own our Datadog and Sentry setup - define SLOs, build dashboards, and keep alerting actionable instead of noisy
  • Lead incident response and run blameless postmortems that actually change how we build
  • Harden our security posture: IAM, secrets management, network boundaries, Cognito auth flows, and vulnerability remediation
  • Protect sensitive health data and support our GDPR and compliance requirements
  • Monitor and optimize AWS spend without compromising reliability
  • Contribute to backend development - APIs, event consumers, data pipelines, and Postgres and DynamoDB performance
  • Participate in architectural discussions and mentor engineers on operational excellence
What we’re looking for
Must-have
  • 5+ years in SRE, DevOps, platform, or backend engineering, with real production ownership
  • Deep AWS experience across serverless and containers - Lambda, ECS Fargate, and debugging both under pressure
  • Strong Terraform skills, including module design and managing state across multiple environments
  • Hands-on experience with event-driven architecture (SQS, SNS, EventBridge) and a healthy respect for its failure modes
  • Solid PostgreSQL: query tuning, indexing, connection management, and zero-downtime migrations
  • Production experience with Datadog or a comparable observability platform (Grafana, New Relic, Honeycomb)
  • Comfortable writing production backend code in [Python / Node.js / Go]
  • Genuine on-call and incident response experience - you've led an incident and written the postmortem
  • Strong security fundamentals: IAM, least privilege, secrets, network isolation, common web vulnerabilities
  • Pragmatic about complexity - you reach for the simplest thing that meets the reliability bar
  • Effective communicator who can explain a technical tradeoff without jargon
Nice-to-have
  • AWS DMS or other data migration and replication tooling
  • Compliance experience (GDPR, SOC 2, ISO 27001)
  • Experience in a B2B SaaS environment or healthcare-related product
What you get
  • Real ownership: a small team, short feedback loops, and no layers of approval between you and production
  • Work that matters: the reliability you build directly affects people reaching for mental health support
  • Free access to the nilo app (incl. family support)
  • A dedicated learning budget for your personal and professional development
  • Work abroad for up to 90 days per year (within the EU)
  • Hybrid working model: 2 days/week from the office, 3 days from home
  • Urban Sports Club membership at a discounted price
  • Equity options: you benefit from any increase in nilo's valuation that you've helped to create
  • Regular team and company events
  • Bring your dog to work: we have 4 office dogs

nilo embraces diversity. We strive to create an inclusive workplace where everyone feels welcome, psychologically and physically safe. All applicants will be considered for employment without regard to different ethnic/racial origins, age groups, religions/ideologies, sexual orientations, gender identities, abilities, socio-economic statuses, educational backgrounds, family arrangements and/or any other characteristic. We would like to encourage you to apply even if the technical requirements cannot be met 100%.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer (SRE / Backend) f/m/d
Senior Site Reliability Engineer (SRE / Backend) f/m/d

nilo • Berlin

Hybrid
EUR 110.000 - 150.000
Real ownership in a small team
Mental health platform access for you/
Free nilo app access (incl. family)
+7
Senior Site Reliability Engineer (SRE / Backend) f/m/d
Senior Site Reliability Engineer (SRE / Backend) f/m/d

DUDE CHEM • Berlin

Hybrid
EUR 110.000 - 170.000
Real ownership
nilo app access
Learning budget
+6
Senior Site Reliability Engineer (m/f/d)
Senior Site Reliability Engineer (m/f/d)

TOPdesk • Kaiserslautern

Vor Ort
EUR 90.000 - 150.000
30 days annual vacation
Remote-friendly options
Health and wellness programs
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

B Capital • Deutschland

Remote
EUR 46.000 - 105.000
Work from anywhere
Flexible paid time off
Mental health support services
+3
Senior Site Reliability Engineer (f/m/d)
Senior Site Reliability Engineer (f/m/d)

Personio • Berlin

Hybrid
EUR 90.000 - 130.000
Competitive reward package including 0
28 days of paid vacation +1 day after
Impact Day
+1
Senior Site Reliability Engineer (m/f/d) at TOPdesk
Senior Site Reliability Engineer (m/f/d) at TOPdesk

TOPdesk • Kaiserslautern

Vor Ort
EUR 90.000 - 120.000
Possibility to work remote
Flexible working hours
Senior Site Reliability Engineer (x/f/m)
Senior Site Reliability Engineer (x/f/m)

Doctolib • Deutschland

Hybrid
EUR 90.000 - 130.000
Deutschlandticket (Germany-wide public
Vacation days (28+1, up to 30)
Work from abroad up to 10 days/year
+8
Staff Site Reliability Engineer (d/f/m)
Staff Site Reliability Engineer (d/f/m)

Devops Academy • München

Hybrid
EUR 70.000 - 90.000
Competitive reward package
28 days of paid vacation
Fully paid Impact Day
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Meyandy LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Pforzheim

Hybrid
EUR 90.000 - 125.000
Hybrid work model