Senior / Staff DevOps & Site Reliability Engineer

Creandum

Baunatal

Hybrid

EUR 90.000 - 130.000

Vollzeit

14 Tage+
Bewerbungsgenerator

A complete application in a minute — tailored resume and cover letter, ready to send.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Flexible remote or hybrid work
Competitive salary and equity
Professional development budget
Greenfield infrastructure project
Impactful work with AI tooling

Zusammenfassung

Creandum is seeking a Senior/Staff DevOps & Site Reliability Engineer to own the infra powering our AI platform. You’ll shape observability, security, and CI/CD, building an AI-enabled development stack where multiple agents work in parallel with human engineers.

You’ll be the first dedicated infra hire, setting standards for scalable, resilient systems and cost-aware cloud architecture across load, team growth, and AI workloads.

Qualifikationen

  • 5+ years in DevOps/SRE or platform engineering with deep AWS expertise.
  • Experience designing infrastructure for AI-native development and IaC practices.
  • Proven track record of building observability stacks and cost-optimized, scalable infra.

Aufgaben

  • Own and evolve AWS infrastructure: containers, networking, security, and cost optimization.
  • Design and implement monitoring and observability with Datadog and related tools.
  • Build proactive systems for incident detection and response.
  • Architect and deploy infra for AI-native development environments.
  • Prepare infra for AI workloads with GPU scheduling and scalable patterns.
  • Create a developer platform where coding agents access data, secrets, and pipelines.
  • Design fast, reliable CI/CD pipelines and deployment workflows for humans and agents.
  • Collaborate with engineering to ensure scalable, future-proof architecture.
  • Establish infrastructure-as-code practices, runbooks, and documentation.

Kenntnisse

DevOps/SRE
AWS
Container orchestration
CI/CD
Observability
Automation
LLM-enabled workflows

Tools

Terraform
Pulumi
GitHub Actions
Datadog
Grafana
Prometheus
CloudWatch

Jobbeschreibung

Full-Time | Remote / Hybrid | Engineering

About the Role

We're scaling fast — and we want to do it without the chaos that usually comes with it.

As our Senior/Staff DevOps & Site Reliability Engineer, you'll own the infrastructure that powers Superscale's AI platform. But this isn't a traditional "keep the lights on" SRE role. You'll be building an infrastructure layer designed for a new kind of engineering team: one where every developer works alongside multiple AI coding agents, and the infra itself is a force multiplier.

You'll be our first dedicated infrastructure hire, which means you get to set the standard — from observability and incident response to CI/CD pipelines and cloud architecture. You'll make sure we scale smoothly as load, team size, and AI workloads grow, and you'll be the counterpart engineers rely on to ship systems that are resilient from day one.

We believe in hiring for breadth and building leverage through AI tooling. We're not growing the team by stacking people in the same roles — we're hiring unique skill sets and amplifying everyone through best-in-class infrastructure and AI-native workflows. You'll be central to making that philosophy real.

Key Responsibilities
  • Own and evolve our AWS infrastructure: containerized services, networking, security, and cost optimization — building toward a setup that scales with both user load and AI workloads
  • Design and implement state-of-the-art monitoring, alerting, and observability with Datadog (no more "is this broken for everyone?" Slack messages — you'll know before anyone asks)
  • Build proactive systems for incident detection and response — shifting the team from reactive firefighting to confident, data-informed operations
  • Architect and deploy infrastructure for AI-native development: cloud-based coding agent environments where multiple agents per developer can build, test, and deploy in parallel
  • Prepare our infrastructure for AI-specific load patterns: bursty GPU/LLM workloads, intelligent request routing, and cost-efficient scaling strategies
  • Create a developer platform that treats coding agents as first-class citizens — giving them access to the same data, tools, secrets, and deployment pipelines that human engineers use
  • Design CI/CD pipelines and deployment workflows that are fast, reliable, and safe — optimized for high-frequency pushes from both humans and agents
  • Partner with the engineering team to build systems that are scaling- and future-proof from the architecture level, not patched after the fact
  • Establish infrastructure-as-code practices, documentation, and runbooks that make the whole team more autonomous
Requirements
  • 5+ years of experience in DevOps, SRE, or platform engineering, with deep hands-on AWS expertise
  • Strong experience with container orchestration (ECS or Kubernetes), infrastructure-as-code (Terraform, Pulumi), and modern CI/CD systems (e.g GitHub Actions)
  • Proven track record of building observability stacks (Datadog, Grafana, Prometheus, CloudWatch, or similar) that actually prevent incidents, not just log them
  • Experience designing infrastructure for service-oriented architectures with relational databases and modern web frontends (Next.js experience is a plus)
  • You understand load balancing, auto-scaling, and cost optimization at a level where you can make real architectural trade-offs
  • Security-minded: you bake in least-privilege access, secrets management, and network segmentation without making developers hate their lives
  • AI-native working style: you actively use LLMs, coding agents, and automation tools in your own workflow. We're building toward 10x coding agents per developer — you'll be the one making that infrastructure possible
  • Strong communicator who can translate infrastructure decisions into language the product and engineering teams understand
Nice to Have
  • Experience building developer platforms or internal tooling that improved team velocity measurably
  • Background in managing AI/ML infrastructure: GPU scheduling, model serving, LLM gateway/proxy setups
  • Experience at an early-stage startup where you built infra foundations that lasted through 10x growth
  • Contributions to open-source infrastructure or DevOps tooling
What We Offer
  • Competitive salary and equity/stock options in a high-growth AI company
  • Flexible remote or hybrid work arrangement
  • Generous paid time off and company holidays
  • Professional development budget for conferences, courses, and certifications
  • Greenfield opportunity — you're setting the infrastructure standard
  • A team that values horizontal skill over narrow specialization, and is investing in AI tooling and agent infrastructure
  • Direct, visible impact — every engineer and every AI agent on the team will feel the quality of what you build

We are an equal opportunity employer and welcome candidates of all backgrounds.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior / Staff Data & ML Engineer
Senior / Staff Data & ML Engineer

Creandum • Baunatal

Vor Ort
EUR 110.000 - 150.000
Flexible remote / hybrid work
Equity / stock options
Professional development budget
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

CloudFactory Limited • Berlin

Vor Ort
EUR 90.000 - 130.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

CloudFactory • Berlin

Vor Ort
EUR 90.000 - 130.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Meyandy LLC • Berlin

Vor Ort
EUR 90.000 - 130.000
Machine Learning Engineer
Machine Learning Engineer

DUDE CHEM • Berlin

Hybrid
EUR 70.000 - 110.000
Senior Product Designer
Senior Product Designer

Creandum • Baunatal

Vor Ort
EUR 70.000 - 100.000
Competitive salary
Meaningful equity
Remote-friendly
+1
Senior/Staff Platform Engineer (m/f/x)
Senior/Staff Platform Engineer (m/f/x)

Cortea • Berlin

Hybrid
EUR 90.000 - 140.000
Equity
Competitive salary
Autonomy
+1
Senior AI Platform Engineer with LangGrap
Senior AI Platform Engineer with LangGrap

Aether Biomedical • Deutschland

Vor Ort
EUR 90.000 - 130.000
Vacation days
Illness days
Health insurance
+9
Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

Andromeda • Deutschland

Vor Ort
EUR 129.994 - 181.992
Platform Security Engineer
Platform Security Engineer

Superhuman Labs, Inc. • Berlin

Vor Ort
EUR 120.000 - 180.000
Relocation support
Visa assistance
Home office setup
+2