ML Ops Engineer

Anaplan

Greater London

On-site

GBP 90,000 - 150,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Anaplan is seeking an experienced ML Ops Engineer to join our Platform Engineering team to design, scale, and maintain high‑performance MLOps and LLMOps infrastructure for our AI‑infused platform.

You will collaborate with Data Scientists, ML Engineers, and Cloud Infra to streamline model training, deployment, and inference while ensuring GPU utilisation, reliability, and cost efficiency.

Qualifications

  • Hands-on production experience in DevOps, SRE, or Platform Engineering with AI/ML infra.
  • Proven track record deploying, scaling and operationalising ML models/LLMs in cloud-native production.
  • Experience managing compute-intensive GPU infra and HPC environments.
  • Advanced proficiency in Kubernetes, Docker, Helm, KubeFlow, Istio.
  • Hands-on with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.
  • Experience with vLLM, Ray, MLflow, LangChain/LangSmith, DeepSpeed, or Hugging Face TGI.
  • Solid background in AWS / GCP / Azure, Kubecost, GPU cost optimisation.
  • Strong Python/Bash/Go and Linux performance monitoring.

Responsibilities

  • Provision and manage cloud-native AI/ML infra using Kubernetes, Docker and GPU orchestration.
  • Automate core infra with IaC tools (Terraform, Helm, Ansible).
  • Optimize GPU compute, networking and storage for model training and inference.
  • Build and maintain CI/CD and MLOps pipelines for training, packaging and deployment.
  • Deploy LLMs and generative AI workloads with Triton, vLLM, TensorRT-LLM.
  • Enable automated model validation and monitoring for drift and latency.
  • Monitor cloud spend across AWS/GCP/Azure and optimize costs.
  • Implement auto-scaling and spot policies to reduce waste.
  • Establish benchmarking and telemetry for throughput and economics.
  • Implement observability with Prometheus, Grafana, OpenTelemetry, MLflow/Weights & Biases.

Skills

Kubernetes
Docker
Helm
KubeFlow
Istio
Terraform
Ansible
GitHub Actions
ArgoCD
Jenkins
Ray
MLflow
LangChain
LangSmith
DeepSpeed
HuggingFace TGI
AWS / GCP / Azure

Tools

Terraform
Ansible
GitHub Actions
ArgoCD
Jenkins
Ray
MLflow
HuggingFace TGI

Job description

At Anaplan, we are a team of innovators focused on optimizing business decision-making through our leading AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market.

What unites Anaplanners across teams and geographies is our collective commitment to our customers’ success and to our Winning Culture.

Our customers rank among the who’s who in the Fortune 50. Coca-Cola, LinkedIn, Adobe, LVMH and Bayer are just a few of the 2,400+ global companies who rely on our best-in‑class platform.

Our Winning Culture is the engine that drives our teams of innovators. We champion diversity of thought and ideas, we behave like leaders regardless of title, we are committed to achieving ambitious goals, and we love celebratingour wins - big and small.

Supported by operating principles of being strategy‑led, values‑based and disciplined in execution, you’ll be inspired, connected, developed and rewarded here. Everything that makes you unique is welcome; join us and let’s build what’s next - together!

Role Overview

We are seeking a ML Ops Engineer to join our Platform Engineering team at Anaplan. In this role, you will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our cutting‑edge AI-infused scenario planning platform.

You will work closely with Data Scientists, ML Engineers, and Cloud Infrastructure teams to streamline model training, deployment, and inference while ensuring optimal GPU utilisation, reliability, and cost‑efficiency.

Your Impact
  • Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray).
  • Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible.
  • Optimise GPU compute workloads, high-speed networking, and storage for efficient model training and low-latency inference.
  • Build and maintain robust CI/CD and MLOps pipelines for continuous model training, evaluation, packaging, and production deployment.
  • Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM).
  • Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks.
  • Monitor and optimise cloud spend across high-cost GPU/CPU clusters across AWS, GCP, or Azure.
  • Implement auto-scaling strategies, spot instance policies, and dynamic resource allocation to eliminate infrastructure waste.
  • Establish benchmarking and telemetry to track unit economics and throughput for training and serving AI models.
  • Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow.
Your Skills
  • Hands‑on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ML infrastructure.
  • Proven track record of deploying, scaling, and operationalising machine learning models and LLMs in cloud-native production environments.
  • Demonstrated experience managing compute‑intensive GPU infrastructure and high‑performance computing (HPC) environments.
  • Advanced proficiency in Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio).
  • Hands‑on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.
  • Experience with vLLM, Ray, MLflow, LangChain / LangSmith, DeepSpeed, or Hugging Face TGI.
  • Solid background in AWS / GCP / Azure, Kubecost, and GPU cost optimisation techniques.
  • Strong skills in Python, Bash, or Go; deep knowledge of Linux kernel tuning and performance monitoring.
Our Commitment to Diversity, Equity, Inclusionand Belonging (DEIB)

We believe attracting and retaining the best talent and fostering an inclusive culture strengthens our business. DEIB improves our workforce, enhances trust with our partners and customers, and drives business success. Build your career in a place where diversity, equity, inclusion and belonging aren’t just words on paper – this is what drives our innovation, it’s how we connect, and it contributes to what makes us a market leader. We believe in a hiring and working environment where all people are respected and valued, regardless of gender identity or expression, sexual orientation, religion, ethnicity, age, neurodiversity, disability status, citizenship, or any other aspect which makes people unique. We hire you for who you are, and we want you to bring your authentic self to work every day!

We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, perform essential job functions, and receive equitable benefits and all privileges of employment. Please contact us to request accommodation.

C andidate data processed during our recruitment activities is handled in accordance with our Candidate Privacy Notice. This may include the use of artificial intelligence or automated tools to assist our team in evaluating qualifications.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Ops Engineer
ML Ops Engineer

Anaplan Inc • Greater London

Hybrid
GBP 110,000 - 140,000
Platform Lead - ML Ops
Platform Lead - ML Ops

anaplan • Greater London

On-site
GBP 120,000 - 180,000
ML Ops Lead
ML Ops Lead

Anaplan • Greater London

On-site
GBP 120,000 - 180,000
ML Ops Lead
ML Ops Lead

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 180,000
Platform Lead - ML Ops
Platform Lead - ML Ops

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 180,000
Principal Engineer - AI
Principal Engineer - AI

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 150,000
Principal AI Data Scientist
Principal AI Data Scientist

Anaplan • Greater London

On-site
GBP 150,000 - 190,000
Principal AI Engineer
Principal AI Engineer

Anaplan Inc • Greater London

Hybrid
GBP 90,000 - 150,000
Principal AI Engineer
Principal AI Engineer

Anaplan • Greater London

On-site
GBP 120,000 - 180,000
Senior ML Engineer
Senior ML Engineer

Anaplan Inc • Manchester

Hybrid
GBP 90,000 - 130,000