Senior Platform & Reliability Engineer (all genders)

Contabo

United States

Hybrid

USD 104,000 - 150,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Remote or hybrid with flexible hours
Workation across the EU
EGYM Wellpass
Corporate discounts program
30 days vacation + holidays off
Extra day off for social

Job summary

Contabo seeks a Senior Platform & Reliability Engineer to own core infrastructure services, including API gateways, storage, identity, and observability. You will design on-call runbooks, improve distributed tracing and dashboards, and advise across teams on shared infrastructure.

The role emphasizes pragmatic stability and knowledge sharing across a remote-first organization. You will work with Kong, Nginx Ingress, Ceph, Longhorn, Vault, Keycloak, Cloudflare, Prometheus, Grafana, and

Qualifications

  • 7+ years in platform, infrastructure, or SRE roles with end-to-end responsibility.
  • Hands-on storage (Ceph) and Kubernetes persistent storage (Longhorn).
  • Experience with API gateways/ingress (Kong, Nginx Ingress).
  • CDN/edge security, DDoS mitigation, WAF config (Cloudflare).
  • Strong system-design fundamentals and trade-off judgment.
  • Observability experience (Prometheus, Grafana, OpenTelemetry) desirable.
  • Secrets/identity infra (Vault, Keycloak); on-call process drafting a plus.

Responsibilities

  • Design and establish an on-call process with runbooks.
  • Mature observability: tracing, SLOs/SLIs, dashboards across the stack.
  • Advise across teams on shared infrastructure services.
  • Document knowledge and distribute it to prevent single points of failure.
  • Evaluate aging components and decide fixes, replacements, or retirements.

Skills

SRE experience
English fluency
On-call design
Documentation
Cross-team collaboration

Tools

Kong
Nginx Ingress
Ceph
Longhorn
Vault
Keycloak
Cloudflare
Prometheus
Grafana
OpenTelemetry
Kubernetes
Vault/Identity infra

Job description

Your creative field

We are looking for a full-time, permanent Senior Platform & Reliability Engineer (all genders) to start as soon as possible. We live remote-first, but you have the freedom to choose whether you want to work hybrid or completely on-site due to your proximity to one of our locations (Berlin, Cologne, Hamburg, Munich).

As a Senior Platform & Reliability Engineer at Contabo, you take architectural ownership of our shared infrastructure services – the foundation that multiple development teams build on: API gateway and ingress (Kong, Nginx Ingress), persistent storage (Ceph, Longhorn), secrets and identity infrastructure (Vault, Keycloak), edge security and WAF (Cloudflare), and our observability stack. You'll be taking over grown, partially under-documented systems – and that's exactly where the appeal of this role lies: you work your way deep into these systems, identify and remediate known weak points, evaluate aging components with solution-agnostic build-vs-buy reasoning, and decide which legacy pieces get fixed, replaced, or retired.

One of your central mandates: you design and establish an on-call process, including runbooks for platform and infrastructure incidents – where no formal process exists today. In parallel, you mature our observability practices, rolling out distributed tracing, SLOs/SLIs, and meaningful dashboards across the stack, building on our existing tooling with Prometheus, Grafana, Alloy, and OpenTelemetry. In system design reviews, you bring your strong grounding in the fundamentals – load balancing, caching, sharding, replication, consistency trade-offs – and apply them to new and existing services alike.

You won't be managing a team, but you will be the technical authority multiple teams rely on: you advise across teams on shared infrastructure services, document your knowledge consistently, and actively distribute it – so that critical know-how never again depends on a single person. Success in this role means: known risks are resolved, a solid on-call process is up and running, and single points of failure are measurably reduced platform-wide. The position is remote (Germany), with hybrid or on-site work optional; occasional travel for datacenter visits and team offsites is part of the role.

What convinces us

Your personality, paired with:

  • 7+ years of experience in platform, infrastructure, or SRE roles, ideally with end-to-end responsibility for a private cloud or IaaS platform
  • Hands-on production experience with distributed storage systems (Ceph) and Kubernetes persistent storage (Longhorn or comparable)
  • Experience operating API gateways and ingress (Kong, Nginx Ingress, or comparable), including debugging cross-cutting concerns like CORS and rate limiting
  • Experience with CDN/edge security, DDoS mitigation, and WAF configuration (e.g. Cloudflare), as well as firewall-rule design
  • A strong grounding in system design fundamentals (load balancing, caching, sharding/replication, consistency models, message queues) – and the judgment to apply them to real-world trade-offs
  • Experience designing or maturing observability (tracing, metrics, SLOs/SLIs) with tools such as Prometheus, Grafana, and OpenTelemetry is desirable
  • Experience with secrets/identity infrastructure (Vault, Keycloak), messaging systems (NATS), and building on-call processes and incident runbooks from scratch is a plus
  • Composure working with grown, incompletely documented systems – paired with the right mix of pragmatism and perfectionism: fix what's broken, retire what's dead, rather than rewriting everything
  • Genuine enjoyment of acting as the technical go-to person across teams, actively sharing and documenting your knowledge
  • Professional fluency in English (our working language); German language skills, certifications (CKA/CKS, Ceph training), and experience with virtualization platforms (Proxmox, OpenStack) are a plus
Awesome Prospects

At Contabo, we are constantly evolving. Our growth creates space for new ideas, ownership, and real impact for people who want to make a difference.
What defines us is trust, direct communication, and an open feedback culture. We believe the best solutions are built together and that everyone has the opportunity to contribute, grow, and leave their own footprint.

Additionally, we offer:
  • Work the way that fits your life – Remote or hybrid with flexible hours for a healthy work-life balance and the same technical equipment at home as in the office
  • Experience real freedom – Workation across the EU and in our summer office in Mallorca
  • Stay active and healthy – Access to EGYM Wellpass and thousands of fitness and wellness facilities
  • Enjoy exclusive perks – Attractive discounts on many products and services through our corporate benefits program
  • Recharge your energy – 30 days of vacation plus additional days off on Christmas Eve and New Year’s Eve
  • Give back – An extra day off to get involved in social
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Brand Manager (all genders)
Brand Manager (all genders)

Contabo • United States

Hybrid
USD 80,000 - 137,000
Remote or hybrid work options
Flexible hours
EU offices & Mallorca office access
Senior Infrastructure Solutions Engineer arbeitnow planetlabs Remote · 9/26/2026
Senior Infrastructure Solutions Engineer arbeitnow planetlabs Remote · 9/26/2026

Primetime • San Francisco (CA)

Hybrid
USD 88,000 - 110,000
PTO
Wellness program
Home office reimbursement
+4
Weekend Site Reliability Engineer
Weekend Site Reliability Engineer

Sporty Group • United States

On-site
USD 120,000 - 180,000
Remote first
Bonuses (quarterly)
28 days leave
+4
Platform Engineer
Platform Engineer

Jobgether • Kentucky

Remote
USD 110,000 - 140,000
Remote-First Flexibility
Access to premium hardware and software
Generous time off
+3
Senior DevOps Engineer
Senior DevOps Engineer

Adapty.io • United States

Remote
USD 180,000 - 240,000
Flexible remote work
Health insurance
Laptop provided
Cloud Engineer – AWS, IoT & AI in Agile Global Team
Cloud Engineer – AWS, IoT & AI in Agile Global Team

QUNDIS GmbH • Kentucky

Hybrid
USD 80,000 - 100,000
Performance-based compensation
Professional development opportunities
Team events
Senior Cloud Infrastructure Engineer (all genders)
Senior Cloud Infrastructure Engineer (all genders)

Lesson Nine • Austin (TX)

On-site
USD 120,000 - 180,000
30 vacation days
3-month Sabbatical
Family and life situation counseling
+7
Senior Platform & DevOps Engineer
Senior Platform & DevOps Engineer

Berlitz • United States

Remote
EUR 90,000 - 130,000
Platform Engineer
Platform Engineer

finmid.com • United States

Hybrid
USD 150,000 - 210,000
30 days PTO
Equity participation
Home office stipend
+2
Senior Platform Security Engineer (100% Remote within Spain)
Senior Platform Security Engineer (100% Remote within Spain)

United States Digital Space LLC • United States

Hybrid
USD 104,000 - 150,000
Private healthcare
Gym access
English & Spanish classes
+6