Senior ML Ops Engineer

Jobtailor

Deutschland

Vor Ort

EUR 90.000 - 120.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Jobtailor is hiring for a role focused on building and operating ML platforms in production. You will extend and run CI/CD pipelines, model orchestration, and automated training to scale with no manual intervention.

You will own deployment and serving standards, ensuring low latency and high availability across ML services, and build reliable, self-service infrastructure for the organization.

Qualifikationen

  • Experience building ML platforms in production environments.
  • Solid working knowledge of containerization and orchestration (Docker, Kubernetes).
  • Familiarity with ML lifecycle tooling and model serving at scale.

Aufgaben

  • Build and maintain ML infrastructure end-to-end including CI/CD pipelines and training automation.
  • Own model deployment and serving with low latency and high availability.
  • Develop core MLOps capabilities like feature stores and model registries.

Kenntnisse

Python
Docker
Kubernetes
CI/CD
Observability

Tools

Prometheus
Grafana
Datadog

Jobbeschreibung

Responsibilities
  • Build and maintain ML infrastructure end-to-end: Extend and operate the infrastructure that powers every model we ship — including CI/CD pipelines, model orchestration, and automated training pipelines designed to scale reliably without manual intervention.
  • Own model deployment and serving: Help define and evolve the standards and tooling for model serving, ensuring low latency and high availability across our ML services.
  • Develop core MLOps capabilities: Establish and maintain essential infrastructure that functions as reliable, self-service systems for the entire machine learning organization — with a focus on feature stores, model registries, and automated monitoring for performance and data drift.
  • Operationalize infrastructure for the ML team: Collaborate with Operations to enable Kubernetes (k8s) autoscaling and GPU provisioning, turning these into accessible, self-service tools for ML practitioners — including standing up and operating a Kubernetes-based development cluster and taking models from experimentation to GPU-backed production.
  • Improve platform reliability and performance: Partner with Operations to design resilient monitoring using advanced observability tooling. Define service-level objectives and implement automation to reduce manual interventions and improve system reliability.
  • Empower Data Scientists through standardized, optimized workflows: Amplify the impact of the ML team by building clear, well-supported "golden paths" — standardized workflows that streamline the model development lifecycle and let Data Scientists focus on modeling while you handle the infrastructure.
Requirements
  • Experience building and operating ML platforms in production environments.
  • Solid working knowledge of containerization and orchestration (Docker, Kubernetes), Linux internals, and model serving at scale.
  • Familiarity with ML lifecycle tooling, including orchestration frameworks, feature stores, model registries, and drift or performance monitoring.
  • Experience owning production systems: defining service-level objectives (SLOs), building observability (for example, using tools such as Prometheus, Grafana, or Datadog), participating in incident response, and diagnosing large-scale failures systematically. You look for opportunities to automate repetitive work rather than absorb it.
  • Comfort writing production-quality code in Python or a comparable language.
  • Experience modernizing production infrastructure with attention to reliability, risk, and cost — including thoughtful sequencing of work to maintain availability and continuity for live systems.
  • The ability to take ownership of technical outcomes, advocate for decisions using data, and communicate clearly in writing and in person — to both technical and non-technical audiences.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior MLOps Engineer
Senior MLOps Engineer

Jobtailor • Deutschland

Remote
EUR 90.000 - 130.000
Senior Machine Learning Ops Engineer
Senior Machine Learning Ops Engineer

Meyandy LLC • Berlin

Vor Ort
EUR 90.000 - 130.000
Principal ML Platform Engineer Europe
Principal ML Platform Engineer Europe

SLAMcore • Deutschland

Vor Ort
EUR 70.000 - 100.000
Machine Learning Platform Engineer I
Machine Learning Platform Engineer I

Mollie • Deutschland

Hybrid
EUR 85.000 - 120.000
Senior MLOps Engineer
Senior MLOps Engineer

RAIVA Technologies UG (haftungsbeschränkt) • Potsdam

Vor Ort
EUR 70.000 - 90.000
Senior ML Ops Engineer
Senior ML Ops Engineer

Centric Software • Berlin

Vor Ort
EUR 90.000 - 140.000
Python Software Engineer – Machine Learning Systems
Python Software Engineer – Machine Learning Systems

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
Machine Learning Systems & Infrastructure Engineer
Machine Learning Systems & Infrastructure Engineer

SpAItial • München

Vor Ort
EUR 70.000 - 90.000
Lead AI Application Engineer – Infrastructure, LLMOps
Lead AI Application Engineer – Infrastructure, LLMOps

Jobtailor • Deutschland

Remote
EUR 120.000 - 180.000
Senior MLOps Engineer (ML Workflows Engineering)
Senior MLOps Engineer (ML Workflows Engineering)

United States Digital Space LLC • Berlin, München

Hybrid
EUR 85.000 - 135.000