Senior Site Reliability Engineer (SRE / Backend)

nilo.health

Berlin

Vor Ort

EUR 110.000 - 140.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Fully remote / in-office flexibility
Equity options
Office dogs in Berlin & Munich
Urban Sports Club membership
Team events
90 days abroad per year

Zusammenfassung

nilo.health is seeking a Senior Site Reliability Engineer to own the infrastructure powering our wellness platform. Our stack runs on AWS with serverless and event-driven services; Terraform is used to manage everything as code.

The role is hands-on, split about 70% infrastructure and 30% backend, with production ownership and incident response responsibilities. You will work with Datadog, SLOs, and security best practices to protect sensitive health data and optimize costs.

Qualifikationen

  • 5+ years in SRE/DevOps with real production ownership.
  • Hands-on with AWS serverless and event-driven architectures.
  • Experience designing resilient, observable systems and incident postmortems.

Aufgaben

  • Own AWS infrastructure end to end (Lambda, ECS Fargate, SQS/SNS, EventBridge, Cognito, DynamoDB, RDS Postgres, and DMS).
  • Develop and maintain Terraform modules with clean state management and review workflows.
  • Build and maintain CI/CD pipelines for web, mobile backends and infrastructure.
  • Design event-driven services with retries, dead-letter queues, idempotency, and graceful degradation.
  • Lead incident response and run blameless postmortems to drive improvements.
  • Harden security posture: IAM, secrets management, network boundaries, vulnerability remediation.
  • Monitor and optimize AWS spend without compromising reliability.
  • Contribute to backend development — APIs, data pipelines, and database performance.
  • Participate in architectural discussions and mentor engineers on operational excellence.

Kenntnisse

AWS
Serverless
Terraform
SRE
Observability
Datadog
PostgreSQL
Incident response
Security basics
Python/Node.js/Go

Tools

Terraform
CI/CD
SQS/SNS/EventBridge

Jobbeschreibung

  • We’re looking for a Senior Site Reliability Engineer to take ownership of the infrastructure behind our platform. We run web applications and two mobile apps entirely on AWS, built around serverless and event-driven services and managed with Terraform
  • This is a hands-on role with real ownership — you’ll set the standards for how we run production rather than inherit someone else’s. The split is roughly 70% infrastructure and reliability work, 30% backend development. Most of your time goes to the platform, but you’ll be comfortable dropping into the application code to debug a slow query, fix a Lambda, or ship an endpoint alongside the product team
  • Because we handle sensitive mental health data, reliability and security aren’t abstract goals here. When someone books a session in a difficult moment, our platform needs to work
  • Own our AWS infrastructure end to end — Lambda, ECS Fargate, SQS, SNS, EventBridge, SES, Cognito, DynamoDB, RDS Postgres, and DMS
  • Manage everything as code in Terraform, with well-designed modules, clean state management, and a solid review workflow
  • Build and maintain CI/CD pipelines with safe rollout and rollback across web, mobile backends, and infrastructure
  • Design our event-driven services for resilience: retries, dead-letter queues, idempotency, graceful degradation
  • Own our Datadog and Sentry setup — define SLOs, build dashboards, and keep alerting actionable instead of noisy
  • Lead incident response and run blameless postmortems that actually change how we build
  • Harden our security posture: IAM, secrets management, network boundaries, Cognito auth flows, and vulnerability remediation
  • Protect sensitive health data and support our GDPR and compliance requirements
  • Monitor and optimize AWS spend without compromising reliability
  • Contribute to backend development — APIs, event consumers, data pipelines, and Postgres and DynamoDB performance
  • Participate in architectural discussions and mentor engineers on operational excellence
Benefits
  • Flexibility to work fully remotely, in-office, or a bit of both
  • Free access to the nilo.health app
  • Enjoy working abroad for up to 90 days per year
  • Urban Sports Club membership at a discounted price
  • Regular team and company events (in compliance with the COVID19 regulations)
  • A strong company culture characterized by team spirit, empathy, respect, trust, courage, and innovation
  • Psychological safety onboarding in your first weeks
  • 30 days of vacation
  • Equity options - you benefit from any increase in the nilo’s valuation that you’ve helped to create
  • Bring your dog to work, we have office dogs in both the Berlin & Munich location
  • Deep AWS experience across serverless and containers — Lambda, ECS Fargate, and debugging both under pressure
  • Effective communicator who can explain a technical tradeoff without jargon
  • Pragmatic about complexity — you reach for the simplest thing that meets the reliability bar
  • 5+ years in SRE, DevOps, platform, or backend engineering, with real production ownership
  • Production experience with Datadog or a comparable observability platform (Grafana, New Relic, Honeycomb)
  • Strong security fundamentals: IAM, least privilege, secrets, network isolation, common web vulnerabilities
  • Solid PostgreSQL: query tuning, indexing, connection management, and zero-downtime migrations
  • Genuine on-call and incident response experience — you’ve led an incident and written the postmortem
  • Strong Terraform skills, including module design and managing state across multiple environments
  • Hands-on experience with event-driven architecture (SQS, SNS, EventBridge) and a healthy respect for its failure modes
  • Comfortable writing production backend code in [Python / Node.js / Go]
  • AWS DMS or other data migration and replication tooling
  • Compliance experience (GDPR, SOC 2, ISO 27001)
  • Experience in a B2B SaaS environment or healthcare-related product
  • We would like to encourage you to apply even if the technical requirements cannot be met 100%
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer (SRE / Backend) f/m/d
Senior Site Reliability Engineer (SRE / Backend) f/m/d

nilo • Berlin

Hybrid
EUR 110.000 - 150.000
Real ownership in a small team
Mental health platform access for you/
Free nilo app access (incl. family)
+7
Senior Site Reliability Engineer, SRE, Backend
Senior Site Reliability Engineer, SRE, Backend

Jobtailor • Berlin

Vor Ort
EUR 90.000 - 130.000
Senior Site Reliability Engineer (m/w/d)
Senior Site Reliability Engineer (m/w/d)

Impower • München

Hybrid
EUR 70.000 - 90.000
Flexible hours
Ownership in projects
Diverse team culture
Senior Platform Engineer (AWS)
Senior Platform Engineer (AWS)

GotPhoto • Berlin

Hybrid
EUR 70.000 - 90.000
Unlimited paid holiday
Education budget
Remote work options
+1
(Senior) Site Reliability Engineer (m/f/d) in Berlin or Konstanz
(Senior) Site Reliability Engineer (m/f/d) in Berlin or Konstanz

United States Digital Space LLC • Deutschland

Hybrid
EUR 60.000 - 90.000
Hybrid working arrangements
Flexible hours
Subsidized sports or yoga courses
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Visa Hunt • Deutschland

Vor Ort
EUR 90.000 - 150.000
Home office budget
Learning & development budget of €1000
Competitive salary
+5
Senior Engineer (Platform)
Senior Engineer (Platform)

Jobgether • Deutschland

Vor Ort
EUR 120.000 - 160.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 130.000
Fully remote
High ownership
Tech exposure
Staff Site Reliability Engineer (x/f/m)
Staff Site Reliability Engineer (x/f/m)

Meyandy LLC • Berlin

Hybrid
EUR 110.000 - 140.000
Senior Site Reliability Engineer (x/f/m)
Senior Site Reliability Engineer (x/f/m)

United States Digital Space LLC • Berlin

Hybrid
EUR 90.000 - 140.000
Deutschlandticket
Vacation days
Health insurance
+7