MLOps Engineer for AI Platform & LLMs

Anaplan Inc

Greater London

Hybrid

GBP 110,000 - 140,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anaplan Inc is seeking a ML Ops Engineer to join our Platform Engineering team. You will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our AI-infused scenario planning platform.

You will collaborate with Data Scientists, ML Engineers, and Cloud Infrastructure teams to streamline model training, deployment, and inference while optimizing GPU utilisation, reliability, and cost-efficiency.

Qualifications

  • Hands-on production experience in DevOps, SRE, or Platform Engineering, with some experience dedicated to AI/ML infrastructure.
  • Proven track record of deploying, scaling, and operationalising machine learning models and LLMs in cloud-native production environments.
  • Demonstrated experience managing compute-intensive GPU infrastructure and HPC environments.
  • Advanced proficiency in Kubernetes (K8s), Docker, Helm, Kubeflow, and service meshes (Istio).
  • Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.
  • Experience with vLLM, Ray, MLflow, LangChain / LangSmith, DeepSpeed, or Hugging Face TGI.
  • Solid background in AWS / GCP / Azure, Kubecost, and GPU cost optimisation techniques.
  • Strong skills in Python, Bash, or Go; deep knowledge of Linux kernel tuning and performance monitoring.

Responsibilities

  • Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray).
  • Automate core platform infrastructure using IaC tools like Terraform, Helm, and Ansible.
  • Optimise GPU compute workloads, high-speed networking, and storage for efficient model training and low-latency inference.
  • Build and maintain robust CI/CD and MLOps pipelines for continuous model training, evaluation, packaging, and production deployment.
  • Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM).
  • Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks.
  • Monitor and optimise cloud spend across high-cost GPU/CPU clusters across AWS, GCP, or Azure.
  • Implement auto-scaling strategies, spot instance policies, and dynamic resource allocation to eliminate infrastructure waste.
  • Establish benchmarking and telemetry to track unit economics and throughput for training and serving AI models.
  • Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow.

Skills

DevOps/SRE/Platform Engineering
ML model deployment
GPU infrastructure
Kubernetes
Docker
Terraform
Ansible
GitHub Actions
ArgoCD/Jenkins
vLLM/Ray/MLflow
LangChain/LangSmith/DeepSpeed
AWS/GCP/Azure

Tools

Kubernetes
Docker
Helm
Kubeflow
Istio
Terraform
Ansible
GitHub Actions
ArgoCD
Jenkins

Job description

Anaplan Inc is seeking a ML Ops Engineer to join our Platform Engineering team. You will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our AI-infused scenario planning platform.

You will collaborate with Data Scientists, ML Engineers, and Cloud Infrastructure teams to streamline model training, deployment, and inference while optimizing GPU utilisation, reliability, and cost-efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Ops Engineer: AI Platform & GPU Infra
ML Ops Engineer: AI Platform & GPU Infra

Anaplan • Greater London

On-site
GBP 90,000 - 150,000
Platform Lead: Scalable MLOps & GenAI Deployments
Platform Lead: Scalable MLOps & GenAI Deployments

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 180,000
ML Ops Lead: GenAI Deployment & FinOps Leader
ML Ops Lead: GenAI Deployment & FinOps Leader

Anaplan • Greater London

On-site
GBP 120,000 - 180,000
ML Ops Engineer
ML Ops Engineer

Anaplan Inc • Greater London

Hybrid
GBP 110,000 - 140,000
GenAI MLOps Lead — Scale, Deploy & Optimize AI
GenAI MLOps Lead — Scale, Deploy & Optimize AI

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 180,000
Platform Lead, ML Ops & GenAI Deployments
Platform Lead, ML Ops & GenAI Deployments

anaplan • Greater London

On-site
GBP 120,000 - 180,000
ML Ops Engineer
ML Ops Engineer

Anaplan • Greater London

On-site
GBP 90,000 - 150,000
ML Ops Lead
ML Ops Lead

Anaplan • Greater London

On-site
GBP 120,000 - 180,000
Platform Lead - ML Ops
Platform Lead - ML Ops

anaplan • Greater London

On-site
GBP 120,000 - 180,000
Platform Lead - ML Ops
Platform Lead - ML Ops

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 180,000