ML Ops Engineer

cmcmarkets

Greater London

On-site

GBP 90,000 - 130,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

CMC Markets in London is seeking an experienced ML Ops Engineer to build and operate platform capabilities that move models from experimentation to production. You will own automation, deployment, observability and controls across the ML lifecycle, partnering with researchers, software engineers, and product teams.

This hands-on role focuses on making ML systems reproducible, scalable, secure and dependable—from packaging and release through serving, monitoring, retraining and incident response.

Qualifications

  • 3-7 years' professional experience in MLOps, ML platform engineering, ML infrastructure, backend engineering, DevOps or SRE.
  • Strong production Python skills, including clean APIs, testing, performance awareness and maintainable services.
  • Experience deploying, serving and operating machine-learning models in production environments.
  • Practical understanding of the ML lifecycle, including training, validation, inference, model release, monitoring and retraining.
  • Experience designing CI/CD workflows and release processes for ML or other production software systems.
  • Hands-on experience with at least one workflow or orchestration system used for ML training, validation or deployment.
  • Comfort working with cloud infrastructure, containers, infrastructure as code and service networking.
  • Strong understanding of observability, monitoring, alerting, incident response and common failure modes in ML systems.
  • Ability to reason about system design, reliability and operational trade-offs-not just individual tools.
  • Clear communication skills and the ability to work effectively with research, engineering, platform, security and product teams.

Responsibilities

  • Build repeatable workflows for model training, validation, promotion, deployment and retraining.
  • Productionise models through packaging, versioning, model registry integration, deployment automation and safe rollback.
  • Design CI/CD pipelines for ML systems, including automated testing, validation, release controls and environment promotion.
  • Manage experiment tracking, model metadata and reproducibility across research and production.
  • Build reusable tooling and platform capabilities that support multiple models and engineering teams.
  • Deploy and operate batch and online inference services in containerised cloud environments.
  • Define and meet availability, latency, throughput and recovery objectives for ML services.
  • Monitor service health, infrastructure, data-quality signals, data drift, prediction drift and model performance decay.
  • Establish dashboards, alerting and operational runbooks so failures are detected and resolved quickly.
  • Support automated or controlled retraining, model promotion, rollback and model retirement.
  • Debug production issues across model, application, infrastructure and critical data-dependency layers.
  • Improve system robustness, scalability and cost efficiency through automation, observability and infrastructure as code.
  • Write production-grade Python for long-running services, deployment tooling and ML workflows.
  • Establish testing, validation, release and incident-management practices for ML systems.
  • Collaborate with platform, security and data engineering teams on reliable model inputs, access controls, secrets, resilience and compliance.
  • Make explicit trade-offs between research flexibility, delivery speed, operational risk and production stability.

Skills

MLOps
Python
CI/CD
Cloud
Containers
Observability
Model serving
Model registry
Experiment tracking

Tools

Docker
Kubernetes

Job description

ML Ops Engineer

London

We're hiring an ML Ops Engineer to build and operate the platform capabilities that take machine-learning models from experimentation into reliable production services.

You'll own the automation, deployment, observability and operational controls around the ML lifecycle, working closely with research engineers, software engineers, platform teams and product teams.

This is not a research role. It is a hands-on engineering role focused on making ML systems reproducible, scalable, secure and dependable, from model packaging and release through to serving, monitoring, retraining and incident response.

What you'll work on
ML lifecycle and platform engineering
  • Build repeatable workflows for model training, validation, promotion, deployment and retraining.
  • Productionise models through packaging, versioning, model registry integration, deployment automation and safe rollback.
  • Design CI/CD pipelines for ML systems, including automated testing, validation, release controls and environment promotion.
  • Manage experiment tracking, model metadata and reproducibility across research and production.
  • Build reusable tooling and platform capabilities that support multiple models and engineering teams.
Model serving and observability
  • Deploy and operate batch and online inference services in containerised cloud environments.
  • Define and meet availability, latency, throughput and recovery objectives for ML services.
  • Monitor service health, infrastructure, data-quality signals, data drift, prediction drift and model performance decay.
  • Establish dashboards, alerting and operational runbooks so failures are detected and resolved quickly.
  • Support automated or controlled retraining, model promotion, rollback and model retirement.
  • Debug production issues across model, application, infrastructure and critical data-dependency layers.
Reliability, security and engineering quality
  • Improve system robustness, scalability and cost efficiency through automation, observability and infrastructure as code.
  • Write production-grade Python for long-running services, deployment tooling and ML workflows.
  • Establish testing, validation, release and incident-management practices for ML systems.
  • Collaborate with platform, security and data engineering teams on reliable model inputs, access controls, secrets, resilience and compliance.
  • Make explicit trade-offs between research flexibility, delivery speed, operational risk and production stability.
Additional responsibilities
  • Maintain personal/professional development to meet the changing demands of the role, including all relevant regulatory and legislative training
  • When dealing with all customers, clients or colleagues ensure that we provide a clear, fair and consistent high quality service that presents a professional and positive image of CMC Markets
  • Take all reasonable steps to ensure appropriate confidentiality
  • Undertake such other duties, training and/or hours of work as may be reasonably required and which are consistent with the general level of responsibility of this role
KEY SKILLS AND EXPERIENCE
  • 3-7 years' professional experience in MLOps, ML platform engineering, ML infrastructure, backend engineering, DevOps or SRE.
  • Strong production Python skills, including clean APIs, testing, performance awareness and maintainable services.
  • Experience deploying, serving and operating machine-learning models in production environments.
  • Practical understanding of the ML lifecycle, including training, validation, inference, model release, monitoring and retraining.
  • Experience designing CI/CD workflows and release processes for ML or other production software systems.
  • Hands-on experience with at least one workflow or orchestration system used for ML training, validation or deployment.
  • Comfort working with cloud infrastructure, containers, infrastructure as code and service networking.
  • Strong understanding of observability, monitoring, alerting, incident response and common failure modes in ML systems.
  • Ability to reason about system design, reliability and operational trade-offs-not just individual tools.
  • Clear communication skills and the ability to work effectively with research, engineering, platform, security and product teams.
Nice to have
  • Prior ownership of model monitoring, drift detection or automated retraining.
  • Familiarity with model registries, feature stores and offline/online feature-consistency challenges.
  • Experience supporting multiple models, services or teams on a shared ML platform.
  • Exposure to regulated or high-reliability production environments.
  • Experience with PyTorch or similar ML frameworks and model-serving technologies.
Technology environment
  • Language: Python
  • ML tooling: PyTorch or similar frameworks, experiment tracking and model registries
  • Workflow orchestration: ML workflows for training, validation, deployment and retraining
  • Deployment: Containers, model-serving frameworks and infrastructure as code
  • Observability: Metrics, logging, tracing, alerting and mo
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Ops Engineer
ML Ops Engineer

CMC MARKETS PLC • City of Westminster

On-site
GBP 70,000 - 100,000
ML Ops Engineer
ML Ops Engineer

CMC Markets • Greater London

Hybrid
GBP 90,000 - 130,000
ML Ops Engineer: Production ML Platform & Reliability
ML Ops Engineer: Production ML Platform & Reliability

cmcmarkets • Greater London

On-site
GBP 90,000 - 130,000
ML Ops Engineer: Production-Ready ML Platform
ML Ops Engineer: Production-Ready ML Platform

CMC Markets • Greater London

Hybrid
GBP 90,000 - 130,000
Machine Learning Operations Engineer
Machine Learning Operations Engineer

Pharmacy2U Ltd • Leeds

Hybrid
GBP 70,000 - 100,000
Pension plan
Sick pay
Long-service awards
+5
MLOps manager
MLOps manager

Uniting Ambition • Greater London

Hybrid
GBP 110,000 - 150,000
ML Platform Engineer: Production ML & Observability
ML Platform Engineer: Production ML & Observability

CMC MARKETS PLC • City of Westminster

On-site
GBP 70,000 - 100,000
MLOps Engineer: Build Production ML Platform
MLOps Engineer: Build Production ML Platform

SMG • City of Westminster

Hybrid
GBP 90,000 - 130,000
Discretionary bonus
Wellbeing fund
Headspace subscription
+2
ML Engineer
ML Engineer

Harnham - Data & Analytics Recruitment • City Of London

On-site
GBP 66,000 - 111,000
Principal Machine Learning Engineer
Principal Machine Learning Engineer

NLP PEOPLE • Greater London

On-site
GBP 80,000 - 120,000