ML Ops Engineer: AI Platform & GPU Infra

Anaplan

Greater London

On-site

GBP 90,000 - 150,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anaplan is seeking an experienced ML Ops Engineer to join our Platform Engineering team to design, scale, and maintain high‑performance MLOps and LLMOps infrastructure for our AI‑infused platform.

You will collaborate with Data Scientists, ML Engineers, and Cloud Infra to streamline model training, deployment, and inference while ensuring GPU utilisation, reliability, and cost efficiency.

Qualifications

  • Hands-on production experience in DevOps, SRE, or Platform Engineering with AI/ML infra.
  • Proven track record deploying, scaling and operationalising ML models/LLMs in cloud-native production.
  • Experience managing compute-intensive GPU infra and HPC environments.
  • Advanced proficiency in Kubernetes, Docker, Helm, KubeFlow, Istio.
  • Hands-on with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.
  • Experience with vLLM, Ray, MLflow, LangChain/LangSmith, DeepSpeed, or Hugging Face TGI.
  • Solid background in AWS / GCP / Azure, Kubecost, GPU cost optimisation.
  • Strong Python/Bash/Go and Linux performance monitoring.

Responsibilities

  • Provision and manage cloud-native AI/ML infra using Kubernetes, Docker and GPU orchestration.
  • Automate core infra with IaC tools (Terraform, Helm, Ansible).
  • Optimize GPU compute, networking and storage for model training and inference.
  • Build and maintain CI/CD and MLOps pipelines for training, packaging and deployment.
  • Deploy LLMs and generative AI workloads with Triton, vLLM, TensorRT-LLM.
  • Enable automated model validation and monitoring for drift and latency.
  • Monitor cloud spend across AWS/GCP/Azure and optimize costs.
  • Implement auto-scaling and spot policies to reduce waste.
  • Establish benchmarking and telemetry for throughput and economics.
  • Implement observability with Prometheus, Grafana, OpenTelemetry, MLflow/Weights & Biases.

Skills

Kubernetes
Docker
Helm
KubeFlow
Istio
Terraform
Ansible
GitHub Actions
ArgoCD
Jenkins
Ray
MLflow
LangChain
LangSmith
DeepSpeed
HuggingFace TGI
AWS / GCP / Azure

Tools

Terraform
Ansible
GitHub Actions
ArgoCD
Jenkins
Ray
MLflow
HuggingFace TGI

Job description

Anaplan is seeking an experienced ML Ops Engineer to join our Platform Engineering team to design, scale, and maintain high‑performance MLOps and LLMOps infrastructure for our AI‑infused platform.

You will collaborate with Data Scientists, ML Engineers, and Cloud Infra to streamline model training, deployment, and inference while ensuring GPU utilisation, reliability, and cost efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Engineer for AI Platform & LLMs
MLOps Engineer for AI Platform & LLMs

Anaplan Inc • Greater London

Hybrid
GBP 110,000 - 140,000
Platform Lead: Scalable MLOps & GenAI Deployments
Platform Lead: Scalable MLOps & GenAI Deployments

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 180,000
ML Ops Lead: GenAI Deployment & FinOps Leader
ML Ops Lead: GenAI Deployment & FinOps Leader

Anaplan • Greater London

On-site
GBP 120,000 - 180,000
Platform Lead, ML Ops & GenAI Deployments
Platform Lead, ML Ops & GenAI Deployments

anaplan • Greater London

On-site
GBP 120,000 - 180,000
GenAI MLOps Lead — Scale, Deploy & Optimize AI
GenAI MLOps Lead — Scale, Deploy & Optimize AI

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 180,000
ML Ops Engineer
ML Ops Engineer

Anaplan Inc • Greater London

Hybrid
GBP 110,000 - 140,000
ML Ops Engineer
ML Ops Engineer

Anaplan • Greater London

On-site
GBP 90,000 - 150,000
ML Ops Lead
ML Ops Lead

Anaplan • Greater London

On-site
GBP 120,000 - 180,000
Platform Lead - ML Ops
Platform Lead - ML Ops

anaplan • Greater London

On-site
GBP 120,000 - 180,000
Platform Lead - ML Ops
Platform Lead - ML Ops

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 180,000