Senior ML Ops Engineer

Boehringer Ingelheim

Greater London

Hybrid

GBP 75,000 - 95,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Boehringer Ingelheim is seeking a Senior MLOps Engineer to ensure AI Accelerator's models transition from development to production reliably. This hybrid role in London will have approximately 3 days in the office each week.

The successful candidate will take operational ownership of models, manage deployment, monitor performance, and ensure MLOps standards are upheld. A Master's degree in a relevant field is required, with a PhD preferred. You should have solid hands-on experience in ML workflows and familiarity with distributed training frameworks.

Qualifications

  • Experience with large foundation model training and multimodal data.
  • Solid hands-on experience operating ML in production environments.
  • Familiarity with distributed training frameworks and CI/CD tooling.

Responsibilities

  • Ensure effective usage of experiment tracking and model registry systems.
  • Manage deployment, monitoring, retraining of models end-to-end.
  • Participate in model handovers, reviewing documentation before acceptance.

Skills

MLOps standards and practices
Machine Learning workflows
Distributed training frameworks
Experiment tracking systems
CI/CD tooling for ML workloads
Cloud infrastructure for ML
Collaboration with research teams

Education

MSc in Machine Learning, Computer Science, or Software Engineering
PhD preferred or equivalent industry experience

Tools

PyTorch
MLflow
Terraform

Job description

We are looking for a Senior MLOps Engineer to join AI Enablement and play a central role in ensuring that the AI Accelerator’s models move from development to production reliably and keep performing. This is a hands‑on operational role with real stakes. The models you deploy and manage will be used to make decisions about which indications to pursue, in which patient population and against which target. When your systems work well, science moves faster and portfolio decision‑making gets better.

You will take full operational ownership of shipped models, managing deployment, monitoring, retraining and lifecycle end‑to‑end. You will make sure that the IT‑provisioned experiment tracking and model registry systems are used effectively, that training and fine‑tuning runs are consistently and correctly logged and that model artefacts are registered with full provenance from data through to prediction. You will work closely with ML engineers at model handover, reviewing documentation and signing off before accepting operational ownership.

This role is for someone who takes pride in operational excellence and who understands that the AI Accelerator models can only realise their impact on the portfolio if they are deployed and performing reliably in production.

Location: London – this is a hybrid role with approximately 3 days a week in the office.

Key Responsibilities
  • Ensure experiment tracking and model registry systems are used effectively across the AI Accelerator with consistent and correct logging of training and fine‑tuning runs, and model artefacts registered with full provenance.
  • Configure, run and troubleshoot distributed training and fine‑tuning jobs, ensuring efficient use of available compute and resolving job‑level failures.
  • Participate in structured model handovers with ML engineers, reviewing and signing off documentation before accepting full model operational ownership of shipped models.
  • Deploy, monitor and manage model serving endpoints, making technical decisions about serving configurations to meet performance requirements of downstream users.
  • Take full operational ownership of models in production, managing monitoring, retraining and lifecycle end‑to‑end.
  • Uphold MLOps standards and practices across the AI Accelerator, contributing to their evolution based on operational experience and keeping teams current with relevant advances in MLOps tooling.
Required Qualifications
  • MSc in Machine Learning, Computer Science, Software Engineering or a related technical field; PhD preferred or equivalent industry experience.
  • Solid hands‑on experience operating ML training and serving workflows in production environments.
  • Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP or Ray Train.
  • Experience operating experiment tracking systems and model registry systems such as MLflow, Weights & Biases or equivalent.
  • Familiarity with CI/CD tooling for ML workflows e.g. cloud‑native pipeline services, GitHub Actions or equivalent.
  • Solid understanding of cloud infrastructure for ML (compute, storage, networking) sufficient to specify requirements clearly and diagnose infrastructure‑related issues.
  • Awareness of large model training characteristics including memory footprint, compute scaling and parallelisation strategies.
  • Familiarity with infrastructure‑as‑code tooling such as Terraform or cloud‑native equivalents.
  • Experience working closely with research and ML engineering teams as a platform operator.
  • Experience operating ML infrastructure for large foundation model training.
  • Familiarity with biomedical AI workloads, such as training foundation models on large‑scale multimodal data.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff MLOps Engineer
Senior Staff MLOps Engineer

Boehringer Ingelheim GmbH • Greater London

Hybrid
GBP 110,000 - 150,000
Top Employer UK
Machine Learning Researcher
Machine Learning Researcher

KEMIO Consulting • Greater London

Hybrid
GBP 75,000 - 120,000
MLOps manager
MLOps manager

Uniting Ambition • Greater London

Hybrid
GBP 110,000 - 150,000
Senior MLOps Engineer - Hybrid London
Senior MLOps Engineer - Hybrid London

Boehringer Ingelheim • Greater London

Hybrid
GBP 75,000 - 95,000
Senior MLOps Engineer
Senior MLOps Engineer

Searchability • Brentford

On-site
GBP 86,000 - 105,000
MLOps Engineer – AI Infrastructure & Deployment
MLOps Engineer – AI Infrastructure & Deployment

Talenzon group • Greater London

On-site
GBP 70,000 - 90,000
MLOPs Engineer with Azure
MLOPs Engineer with Azure

Gazelle Global • Wokingham

On-site
GBP 70,000 - 110,000
Senior Cloud Engineer
Senior Cloud Engineer

Boehringer Ingelheim GmbH • Greater London

Hybrid
GBP 90,000 - 120,000
Machine Learning Operations Engineer
Machine Learning Operations Engineer

Pharmacy2U Ltd • Leeds

Hybrid
GBP 70,000 - 100,000
Pension plan
Sick pay
Long-service awards
+5
MLOPS Lead
MLOPS Lead

Candour Solutions LTD • York and North Yorkshire

On-site
GBP 70,000 - 90,000