Platform Lead - ML Ops

Anaplan Inc

Greater London

Hybrid

GBP 120,000 - 180,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anaplan Inc. seeks a Platform Lead to spearhead ML Ops infrastructure, cost-optimisation, and deployment strategies. You will manage a talented DevOps team while remaining deeply technical and hands-on, building scalable platforms for ML and GenAI models with financial accountability.

You will lead infrastructure strategy, automate provisioning, and deploy LLMs with a focus on reliability, security, and cost visibility across cloud environments.

Qualifications

  • Proven experience building and operating ML/GenAI platforms in production.
  • Demonstrated leadership of engineering teams delivering reliable ML systems.
  • Strong track record reducing cloud spend on large AI clusters.
  • Experience with CI/CD, IaC, and observability in cloud environments.

Responsibilities

  • Lead a DevOps team focusing on ML Ops and GenAI deployment.
  • Define infra roadmap for AI/ML workloads and automate provisioning with IaC.
  • Architect and maintain robust MLOps/LLMOps pipelines and CI/CD frameworks.
  • Deploy LLMs in production with high availability and low latency.
  • Establish FinOps to forecast and manage AI infrastructure spend.
  • Implement auto-scaling, spot instances, and cost-visibility into unit economics.
  • Set up 24/7 observability and incident response for AI infra.
  • Enforce data governance, security, and compliance policies.

Skills

ML Ops
Leadership
GenAI deployment
Cloud cost optimization
Kubernetes
Terraform
Jenkins/GitHub Actions
Python/Bash/Go
LLMs
CI/CD

Tools

MLflow
Kubeflow
LangSmith
Phoenix
Kubecost
Cloudability
Triton Inference Server
Docker
Git

Job description

At Anaplan, we are a team of innovators focused on optimizing business decision-making through our leading AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market.

What unites Anaplanners across teams and geographies is our collective commitment to our customers’ success and to our Winning Culture.

Our customers rank among the who’s who in the Fortune 50. Coca-Cola, LinkedIn, Adobe, LVMH and Bayer are just a few of the 2,400+ global companies who rely on our best-in-class platform.

Our Winning Culture is the engine that drives our teams of innovators. We champion diversity of thought and ideas, we behave like leaders regardless of title, we are committed to achieving ambitious goals, and we love celebratingour wins - big and small.

Supported by operating principles of being strategy-led, values -based and disciplined in execution, you’ll be inspired, connected, developed and rewarded here. Everything that makes you unique is welcome; join us and let’s build what’s next - together!

Role Overview

We are seeking a Platform Lead to spearhead our ML Ops infrastructure, cost-optimisation, and deployment strategies. You will manage a talented team of DevOps engineers while remaining deeply technical and hands-on. Your primary mission is to build, scale, and secure the foundational platforms for our machine learning (ML) and generative AI (GenAI) models while maintaining financial accountability.

Your Impact
  • Team Leadership & Collaboration: Lead and manage a dedicated DevOps team, mentoring both junior and senior engineers while collaborating closely with Data Science and Engineering leaders.
  • Infrastructure Strategy & Automation: Define the infrastructure roadmap for AI/ML workloads and automate provisioning across cloud environments using Infrastructure as Code (IaC).
  • MLOps & LLMOps Engineering: Architect, maintain, and optimise robust MLOps/LLMOps pipelines and CI/CD frameworks for continuous model deployment.
  • GenAI Production Deployment: Deploy Large Language Models (LLMs) into production environments, ensuring high availability, low latency, and optimal performance for GenAI applications.
  • FinOps & Budget Management: Establish FinOps frameworks to track, allocate, and forecast AI infrastructure spend, managing high-cost GPU/CPU cloud budgets.
  • Resource Efficiency & Unit Economics: Implement auto-scaling, spot instances, and down-scaling policies to eliminate waste, while providing full visibility into the unit economics of training and serving LLM models.
  • Observability & Incident Response: Establish 24/7 incident response, telemetry, and observability metrics to monitor system performance, model drift, and data pipelines.
  • Data Governance & Security: Enforce strict data governance, platform security, and compliance protocols across all AI/ML infrastructure.
Your Skills
  • Extensive production experience deploying and supporting ML systems.
  • Proven track record of leading engineering teams.
  • Demonstrated experience with Generative AI and LLM deployment patterns.
  • A proven history of reducing cloud spend on large-scale AI clusters.
  • Experience with tools like MLflow, Kubeflow, LangSmith, or Phoenix.
  • Expertise in AWS/GCP/Azure cost tools, Kubecost, or Cloudability.
  • Extensive background of Kubernetes (K8s), Docker, and service meshes.
  • Expert knowledge of Terraform, Ansible, Jenkins, or GitHub Actions.
  • Proficient in Python, Bash, or Go.
  • Familiarity with Triton Inference Server, vLLM, or Hugging Face TGI.
Our Commitment to Diversity, Equity, Inclusionand Belonging (DEIB)

We believe attracting and retaining the best talent and fostering an inclusive culture strengthens our business. DEIB improves our workforce, enhances trust with our partners and customers, and drives business success. Build your career in a place where diversity, equity, inclusion and belonging aren’t just words on paper – this is what drives our innovation, it’s how we connect, and it contributes to what makes us a market leader. We believe in a hiring and working environment where all people are respected and valued, regardless of gender identity or expression, sexual orientation, religion, ethnicity, age, neurodiversity, disability status, citizenship, or any other aspect which makes people unique. We hire you for who you are, and we want you to bring your authentic self to work every day!

We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, perform essential job functions, and receive equitable benefits and all privileges of employment. Please contact us to request accommodation.

C andidate data processed during our recruitment activities is handled in accordance with our Candidate Privacy Notice. This may include the use of artificial intelligence or automated tools to assist our team in evaluating qualifications.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Ops Lead
ML Ops Lead

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 180,000
ML Ops Lead
ML Ops Lead

Anaplan • Greater London

On-site
GBP 120,000 - 180,000
Platform Lead - ML Ops
Platform Lead - ML Ops

anaplan • Greater London

On-site
GBP 120,000 - 180,000
ML Ops Engineer
ML Ops Engineer

Anaplan Inc • Greater London

Hybrid
GBP 110,000 - 140,000
ML Ops Engineer
ML Ops Engineer

Anaplan • Greater London

On-site
GBP 90,000 - 150,000
AI Technical Lead
AI Technical Lead

Anaplan • Greater London

On-site
GBP 120,000 - 180,000
AI Technical Lead
AI Technical Lead

Anaplan Inc • Greater London

Hybrid
GBP 110,000 - 170,000
Principal Architect AI
Principal Architect AI

Anaplan • Greater London

On-site
GBP 140,000 - 190,000
Engineering Manager - Platform
Engineering Manager - Platform

Anaplan Inc • Greater London

Hybrid
GBP 110,000 - 140,000
Principal Architect AI
Principal Architect AI

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 180,000