Senior Machine Learning Ops Engineer

Meyandy LLC

Berlin

Vor Ort

EUR 90.000 - 130.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

KAYAK in Berlin is seeking a Senior ML Ops Engineer to bridge data science and production engineering. You will build and maintain scalable ML infrastructure, automate training and deployment pipelines, and ensure reliable, reproducible models at scale.

You will collaborate with Data Scientists, ML Engineers, and Operations to transform experimental code into robust services. The role requires commuting to the Berlin office 3 times a week.

Qualifikationen

  • Experience building and operating ML platforms in production environments.
  • Strong knowledge of containerization and orchestration (Docker, Kubernetes) and Linux internals.
  • Familiarity with ML lifecycle tooling, including feature stores and model registries.
  • Experience owning production systems: SLOs, observability, and incident response.
  • Proactive, automation-focused mindset with production-grade code ability.

Aufgaben

  • Build and maintain ML infrastructure end-to-end.
  • Own model deployment, serving, and low-latency ML services.
  • Develop ML lifecycle tooling and automated monitoring.
  • Operate Kubernetes-based infra and GPU provisioning.
  • Improve platform reliability with advanced observability.
  • Create golden-path workflows to streamline ML development.

Kenntnisse

ML Ops
Docker
Kubernetes
Observability

Tools

Prometheus
Grafana
Datadog
Feature stores
Model registries
CI/CD pipelines
GPU provisioning
Kubernetes autoscaling

Jobbeschreibung

Senior Machine Learning Ops Engineer — Kayak, Berlin Office

KAYAK, part of Booking Holdings (NASDAQ: BKNG), is a leading travel search engine. With billions of queries across our platforms, we help people find their perfect flight, stay, rental car and vacation package. We're also transforming business travel with a new corporate travel solution, KAYAK for Business. As an employee of KAYAK, you will be part of a travel company that operates a portfolio of global metasearch brands including momondo, Cheapflights and HotelsCombined, among others. From start-up to industry leader, innovation is in our DNA and every employee has an opportunity to make their mark. Our focus is on building the best travel search engine to make it easier for everyone to experience the world. Every machine learning model KAYAK ships depends on reliable, scalable infrastructure to move from experiment to production — and that's exactly what this role makes possible. This is a senior, hands‑on role where you will bridge the gap between data science and production engineering. You will join the Machine Learning Platform team and be responsible for building and maintaining scalable infrastructure & automated pipelines for model training, deployment, and monitoring, ensuring our ML models are reliable, reproducible, and performant. You will work closely with Data Scientists, ML Engineering and Operations teams to transform experimental code into robust, production-ready services at scale. This role requires commuting to the Berlin office 3 times a week.


In this role, you will:



  • Build and maintain ML infrastructure end-to-end: Extend and operate the infrastructure that powers every model we ship — including CI/CD pipelines, model orchestration, and automated training pipelines designed to scale reliably without manual intervention.

  • Own model deployment and serving: Help define and evolve the standards and tooling for model serving, ensuring low latency and high availability across our ML services.

  • Develop core MLOps capabilities: Establish and maintain essential infrastructure that functions as reliable, self-service systems for the entire machine learning organization — with a focus on feature stores, model registries, and automated monitoring for performance and data drift.

  • Operationalize infrastructure for the ML team: Collaborate with Operations to enable Kubernetes (k8s) autoscaling and GPU provisioning, turning these into accessible, self-service tools for ML practitioners — including standing up and operating a Kubernetes-based development cluster and taking models from experimentation to GPU-backed production.

  • Improve platform reliability and performance: Partner with Operations to design resilient monitoring using advanced observability tooling. Define service-level objectives and implement automation to reduce manual interventions and improve system reliability.

  • Empower Data Scientists through standardized, optimized workflows: Amplify the impact of the ML team by building clear, well-supported \"golden paths\" — standardized workflows that streamline the model development lifecycle and let Data Scientists focus on modeling while you handle the infrastructure.


Please apply if you have:



  • Experience building and operating ML platforms in production environments.

  • Solid working knowledge of containerization and orchestration (Docker, Kubernetes), Linux internals, and model serving at scale.

  • Familiarity with ML lifecycle tooling, including orchestration frameworks, feature stores, model registries, and drift or performance monitoring.

  • Experience owning production systems: defining service-level objectives (SLOs), building observability (for example, using tools such as Prometheus, Grafana, or Datadog), participating in incident response, and diagnosing large-scale failures systematically.

  • You look for opportunities to automate repetitive work rather than absorb it. Comfort writing production-quality code in Pyth…

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Machine Learning Ops Engineer
Senior Machine Learning Ops Engineer

KAYAK • Berlin

Hybrid
EUR 90.000 - 140.000
Remote days
Therapy sessions
HeadSpace access
+1
Senior Machine Learning Ops Engineer
Senior Machine Learning Ops Engineer

Momondo • Berlin

Vor Ort
EUR 90.000 - 130.000
Work from (almost) anywhere for up to
20 days per year remote work
Mental health & well-being programs
+4
Senior ML Ops Engineer
Senior ML Ops Engineer

KAYAK • Berlin

Hybrid
EUR 90.000 - 135.000
Work from almost anywhere
Mental health support
HeadSpace subscription
+11
Senior Machine Learning Ops Engineer
Senior Machine Learning Ops Engineer

United States Digital Space LLC • Berlin

Hybrid
EUR 120.000 - 150.000
Work from almost anywhere (up to 20</n
Mental health support
HeadSpace subscription
+5
Senior ML Ops Engineer
Senior ML Ops Engineer

Jobtailor • Deutschland

Vor Ort
EUR 90.000 - 120.000
Manager, Data Platform
Manager, Data Platform

KAYAK • Berlin

Hybrid
EUR 120.000 - 180.000
Work from almost anywhere (up to 20y/d
Mental health & well-being: therapy &头
No meeting Fridays
+12
Data Platform Manager
Data Platform Manager

KAYAK • Berlin

Hybrid
EUR 140.000 - 190.000
Flexible work policy
Time off to recharge
Volunteer Time Off
+3
ML Product Engineer
ML Product Engineer

Meyandy LLC • Heidelberg

Hybrid
EUR 90.000 - 130.000
VSOP equity
30 days paid holiday
Statutory social insurance
+3
Senior MLOps Engineer
Senior MLOps Engineer

Talon.One • Berlin

Vor Ort
EUR 70.000 - 90.000
€1,000 annual learning budget
30 days of annual leave
Home office setup budget
+3
Principal ML Platform Engineer Europe
Principal ML Platform Engineer Europe

SLAMcore • Deutschland

Vor Ort
EUR 70.000 - 100.000