Staff DevOps Engineer

Nexxa.AI

Deutschland

Remote

EUR 110.000 - 180.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Eine Bewerbung wie gemacht für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Nexxa.ai in Germany is seeking a Senior/Staff DevOps Engineer to own and operate the infrastructure that powers AI and industrial systems at scale. You will ensure production ML and data workloads run fast, observable, and resilient—from training clusters to deployment pipelines.

You will partner with AI, data, and product teams to enable reliable production environments, design IaC with Terraform/Pulumi, manage Kubernetes platforms, and drive SLOs/incident response.

Qualifikationen

  • 6+ years in DevOps, SRE, Platform or infrastructure-focused roles.
  • Hands-on with observability stacks: Prometheus, Grafana, Datadog, OpenTelemetry.
  • Automation and tooling using Python, Go or Bash.
  • Ability to independently scope and lead infrastructure projects from design to production.
  • Strong incident management and root-cause analysis skills.

Aufgaben

  • Own and evolve Nexxa’s core infrastructure end-to-end.
  • Design and operate CI/CD pipelines for AI, data and product teams.
  • Build and maintain IaC for reproducible environments across cloud and on-prem/edge.
  • Architect and manage Kubernetes-based platforms for training and inference workloads.
  • Define observability practices: metrics, logging, tracing, alerting across distributed systems.
  • Enforce reliability: SLOs/SLIs, incident response, postmortems, on-call rotations.
  • Consider security and compliance in cloud, secrets, and access control.

Kenntnisse

DevOps
SRE
Platform Engineering
Scripting (Python/Go/Bash)
IaC tooling

Tools

Prometheus
Grafana
Datadog
OpenTelemetry
Terraform
Pulumi
Kubernetes
Snowflake
BigQuery

Jobbeschreibung

Nexxa is building the best AI systems for heavy industries — enabling machines, systems, and operations to think, decide, and act autonomously across manufacturing, large-scale infrastructure, logistics, and legacy environments.

Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry.

About the Role

We’re looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale. You understand what it takes to keep production ML and data workloads fast, observable, and resilient — from GPU-backed training and inference clusters to the pipelines that connect them to real-world industrial environments.

This role is ideal for candidates who want deep infrastructure ownership at a company where uptime, latency, and reliability directly affect physical operations — not just software. You’ll partner closely with AI, data, and product engineering teams to make sure the systems they build can actually run in production, safely and at scale.

What You’ll Do
  • Own and evolve Nexxa’s core infrastructure — compute, networking, storage, and deployment systems — end-to-end
  • Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams
  • Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments
  • Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling
  • Partner with data and AI teams to support the infrastructure behind:
  • Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems
  • Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations
  • Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations
  • Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity
  • Collaborate with engineering leadership to define infrastructure roadmap and platform strategy
  • Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org
Required Qualifications
  • 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles
  • Deep hands-on experience with:
  • Strong track record designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry)
  • Experience supporting ML/AI infrastructure — training clusters, model serving, data pipelines — a strong plus
  • Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling
  • Proven ability to independently scope and lead infrastructure projects from design through production rollout
  • Strong incident management instincts — you can lead through an outage calmly and drive toward root cause
Preferred Qualifications
  • Experience operating infrastructure that bridges cloud and edge/on-prem environments, especially in industrial or manufacturing contexts
  • Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks)
  • Experience with service mesh, zero-trust networking, or compliance frameworks relevant to industrial/critical infrastructure (e.g., SOC 2, IEC 62443)
  • History of building internal developer platforms or self-service infrastructure tooling
  • Experience scaling infrastructure teams or setting technical direction at a Staff level
What Success Looks Like
  • You can owner ambiguous, high-stakes infrastructure problems end-to-end
  • Systems you build stay reliable as usage and scale grow — you design for the next order of magnitude, not just today
  • You bring strong technical judgment on tradeoffs between reliability, cost, and speed
  • You raise the bar for operational rigor and engineering discipline across the team
  • You help define what’s next for the platform, not just execute what’s known
Why Join Nexxa.ai?
  • Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies
  • Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement
  • Professional Growth: Benefit from significant opportunities for career development and advancement
  • Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions

If you’re passionate about building the infrastructure that powers advanced AI solutions in the real world, we’d love to connect.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Applied AI Engineer
Applied AI Engineer

nexxa • München

Vor Ort
EUR 90.000 - 150.000
Senior / Staff DevOps & Site Reliability Engineer
Senior / Staff DevOps & Site Reliability Engineer

Creandum • Baunatal

Vor Ort
EUR 90.000 - 130.000
Flexible remote or hybrid work
Competitive salary and equity
Professional development budget
+2
Senior Infrastructure Support Engineer
Senior Infrastructure Support Engineer

Nscale • Deutschland

Vor Ort
EUR 103.000 - 147.000
Base salary + equity
Remote-first team
Annual reviews and progression plan
+1
Staff Software Engineer, Infrastructure (Cloud)
Staff Software Engineer, Infrastructure (Cloud)

AeroVect • München

Vor Ort
EUR 120.000 - 180.000
Senior Software Engineer
Senior Software Engineer

Trust In SODA • München

Vor Ort
EUR 90.000 - 130.000
Staff Software Engineer, Infrastructure (Cloud)
Staff Software Engineer, Infrastructure (Cloud)

AeroVect • Berlin

Vor Ort
EUR 120.000 - 180.000
Staff Software Engineer, Infrastructure (Cloud)
Staff Software Engineer, Infrastructure (Cloud)

AeroVect • Berlin

Vor Ort
EUR 120.000 - 180.000
Field CTO
Field CTO

Nebius B.V. • Deutschland

Vor Ort
EUR 172.000 - 211.000
Health Insurance
401(k) Plan
Parental Leave
+2
Staff Software Engineer, Infrastructure (Cloud)
Staff Software Engineer, Infrastructure (Cloud)

AeroVect • München

Vor Ort
EUR 120.000 - 180.000
Senior AI Engineer
Senior AI Engineer

NXT • Berlin

Vor Ort
EUR 110.000 - 140.000
Hybrid work model