Senior Machine Learning Ops Engineer

United States Digital Space LLC

Berlin

Vor Ort

EUR 120.000 - 150.000

Vollzeit

Vor 13 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Work from almost anywhere (up to 20</n
Mental health support
HeadSpace subscription
Company-wide week off
No meeting Fridays
Paid parental leave
Travel discounts
Office in Berlin

Zusammenfassung

Booking Holdings is seeking a Senior MLOps Engineer to design and implement scalable ML infrastructure and production lifecycles in Berlin. You will bridge data science and production engineering, building end-to-end pipelines for model training, deployment, and monitoring.

You will collaborate with Data Scientists, ML Engineering and Operations to transform experimental code into robust services at scale, commuting to the Berlin office as required.

Qualifikationen

  • Build and maintain ML infrastructure end-to-end for model training, deployment and monitoring.
  • Own model deployment and serving with low latency and high availability.
  • Develop core MLOps capabilities including feature stores and model registries.
  • Operationalize infrastructure with Kubernetes autoscaling and GPU provisioning.
  • Improve platform reliability and observability with SLOs and automation.
  • Empower Data Scientists with standardized, golden-path workflows.

Aufgaben

  • Build and maintain ML infrastructure end-to-end: extend and operate infra powering every model — CI/CD, model orchestration, automated training pipelines.
  • Own model deployment and serving with low latency and high availability across ML services.
  • Develop core MLOps: feature stores, model registries, automated monitoring for performance and drift.
  • Operationalize infra: Kubernetes autoscaling, GPU provisioning, self-service tooling for ML practitioners.
  • Improve platform reliability: design resilient monitoring and define SLAs to reduce manual interventions.
  • Empower Data Scientists with golden paths to streamline model development lifecycle.

Kenntnisse

Python
CI/CD
Observability
SLOs
Production engineering

Tools

Docker
Kubernetes
Prometheus
Grafana
Datadog

Jobbeschreibung

the company, part of Booking Holdings (NASDAQ: BKNG), is a leading travel search engine. With billions of queries across our platforms, we help people find their perfect flight, stay, rental car and vacation package. We're also transforming business travel with a new corporate travel solution, the company for Business.

As an employee of the company, you will be part of a travel company that operates a portfolio of global metasearch brands including momondo, Cheapflights and HotelsCombined, among others. From start-up to industry leader, innovation is in our DNA and every employee has an opportunity to make their mark. Our focus is on building the best travel search engine to make it easier for everyone to experience the world.

Every machine learning model the company ships depends on reliable, scalable infrastructure to move from experiment to production — and that's exactly what this role makes possible. the company is seeking a Senior MLOps Engineer who will focus on the design and implementation of our machine learning infrastructure and production lifecycle. This is a senior, hands-on role where you will bridge the gap between data science and production engineering.

You will join the Machine Learning Platform team and be responsible for building and maintaining scalable infrastructure & automated pipelines for model training, deployment, and monitoring, ensuring our ML models are reliable, reproducible, and performant. You will work closely with Data Scientists, ML Engineering and Operations teams to transform experimental code into robust, production-ready services at scale.

This role requires commuting to the Berlin office 3 times a week.

In this role, you will:

  • Build and maintain ML infrastructure end-to-end: Extend and operate the infrastructure that powers every model we ship — including CI/CD pipelines, model orchestration, and automated training pipelines designed to scale reliably without manual intervention.
  • Own model deployment and serving: Help define and evolve the standards and tooling for model serving, ensuring low latency and high availability across our ML services.
  • Develop core MLOps capabilities: Establish and maintain essential infrastructure that functions as reliable, self-service systems for the entire machine learning organization — with a focus on feature stores, model registries, and automated monitoring for performance and data drift.
  • Operationalize infrastructure for the ML team: Collaborate with Operations to enable Kubernetes (k8s) autoscaling and GPU provisioning, turning these into accessible, self-service tools for ML practitioners — including standing up and operating a Kubernetes-based development cluster and taking models from experimentation to GPU-backed production.
  • Improve platform reliability and performance: Partner with Operations to design resilient monitoring using advanced observability tooling. Define service-level objectives and implement automation to reduce manual interventions and improve system reliability.
  • Empower Data Scientists through standardized, optimized workflows: Amplify the impact of the ML team by building clear, well-supported "golden paths" — standardized workflows that streamline the model development lifecycle and let Data Scientists focus on modeling while you handle the infrastructure.

Please apply if you have:

  • Experience building and operating ML platforms in production environments.
  • Solid working knowledge of containerization and orchestration (Docker, Kubernetes), Linux internals, and model serving at scale.
  • Familiarity with ML lifecycle tooling, including orchestration frameworks, feature stores, model registries, and drift or performance monitoring.
  • Experience owning production systems: defining service-level objectives (SLOs), building observability (for example, using tools such as Prometheus, Grafana, or Datadog), participating in incident response, and diagnosing large-scale failures systematically. You look for opportunities to automate repetitive work rather than absorb it.
  • Comfort writing production-quality code in Python or a comparable language.
  • Experience modernizing production infrastructure with attention to reliability, risk, and cost — including thoughtful sequencing of work to maintain availability and continuity for live systems.
  • The ability to take ownership of technical outcomes, advocate for decisions using data, and communicate clearly in writing and in person — to both technical and non-technical audiences.

Benefits and Perks

  • Work from (almost) anywhere for up to 20 days per year
  • Focus on mental health and well-being:
  • Company-paid therapy sessions through SpringHealth
  • Company-paid subscription to HeadSpace
  • Company-wide week off a year – the whole team fully recharges (and returns without a pile-up of work!)
  • No meeting Fridays
  • Paid parental leave
  • Paid volunteer time
  • Focus on your career growth:
  • Development Dollars
  • Leadership development
  • Access to thousands of on-demand e-learnings
  • Travel Discounts
  • Employee Resource Groups
  • 6 weeks paid vacation + a day off for your birthday
  • Free lunch 2 days per week
  • Pension plan contributions
  • Public transportation subsidies
  • Bike leasing program
  • Monthly social events, Thursday happy hours, sports teams
  • An awesome office in Friedrichshain, Berlin

InclusionAt the company, we want everyone to have the space to grow, share ideas and do great work. That's why we're focused on hiring the best talent from all walks of life and experiences, supporting them well and making sure no one feels like they have to fit a mold to belong here.

Need any adjustments for the interview, application or on the job? No problem – just give us a heads-up. We've got you.

#LI-AS1

Find more English Speaking Jobs in Germany on Arbeitnow

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Machine Learning Ops Engineer
Senior Machine Learning Ops Engineer

KAYAK • Berlin

Hybrid
EUR 90.000 - 140.000
Remote days
Therapy sessions
HeadSpace access
+1
Senior Machine Learning Ops Engineer
Senior Machine Learning Ops Engineer

Momondo • Berlin

Vor Ort
EUR 90.000 - 130.000
Work from (almost) anywhere for up to
20 days per year remote work
Mental health & well-being programs
+4
Senior Machine Learning Ops Engineer
Senior Machine Learning Ops Engineer

Meyandy LLC • Berlin

Vor Ort
EUR 90.000 - 130.000
Machine Learning Ops Engineer - Personalization (m|w|d)
Machine Learning Ops Engineer - Personalization (m|w|d)

United States Digital Space LLC • Berlin

Vor Ort
EUR 70.000 - 110.000
Machine Learning Ops Engineer - Personalization (m|w|d)
Machine Learning Ops Engineer - Personalization (m|w|d)

DUDE CHEM • Berlin

Vor Ort
EUR 90.000 - 120.000
Free breakfast
Free lunch
Coffee & beverages
+5
Senior MLOps Engineer
Senior MLOps Engineer

Talon.One • Berlin

Vor Ort
EUR 70.000 - 90.000
€1,000 annual learning budget
30 days of annual leave
Home office setup budget
+3
Senior MLOps Engineer (ML Workflows Engineering)
Senior MLOps Engineer (ML Workflows Engineering)

United States Digital Space LLC • Berlin, München

Hybrid
EUR 85.000 - 135.000
AI Platform Engineer (f/m/d)
AI Platform Engineer (f/m/d)

United States Digital Space LLC • München

Hybrid
EUR 90.000 - 130.000
Hybrid work
Competitive compensation
Pension Plan/Bonus
+5
Senior Engineering Manager, K4B
Senior Engineering Manager, K4B

United States Digital Space LLC • Berlin

Vor Ort
EUR 120.000 - 180.000
Travel discounts
6 weeks paid vacation + birthday
Free lunch 2 days per week
+2
Engineering Manager, K4B
Engineering Manager, K4B

United States Digital Space LLC • Berlin

Hybrid
EUR 110.000 - 160.000
Work from almost anywhere for up to 20
Mental health and well-being focus
Company-paid therapy sessions
+10