We are looking for Senior MLOps Engineer at London, UK – 3 days per week Onsite
Role Overview
We are seeking an experienced Senior MLOps Engineer to support the design, implementation, and optimisation of enterprise-scale MLOps platforms on Microsoft Azure. Working closely with Solution and Enterprise Architects, the successful candidate will help build and operate scalable machine learning platforms on Kubernetes, with a focus on model lifecycle management, observability, low-latency inference, platform reliability, and cost efficiency.
Key Responsibilities
- Partner with Architects to design and implement end-to-end MLOps solutions on Azure.
- Build and operate scalable ML platforms using Azure Kubernetes Service (AKS) and cloud-native technologies.
- Develop CI/CD and Continuous Training (CT) pipelines for machine learning workloads.
- Deploy, manage, and optimise ML workloads in Kubernetes environments.
- Implement model serving capabilities that meet high-availability and low-latency requirements.
- Configure autoscaling, traffic management, rollback strategies, and resource governance.
- Manage containerised ML applications using Docker, Kubernetes, Helm, and GitOps practices.
- Implement monitoring and observability across:
- Model performance and drift
- Application performance and platform health
- Infrastructure and operational metrics
Performance & Cost Optimisation
- Optimise cloud infrastructure utilisation and spend for ML workloads.
- Implement efficient compute and scaling strategies across training and inference environments.
- Drive FinOps practices, cost visibility, and resource right-sizing.
- Improve platform performance, reliability, throughput, and latency.
Required Skills & Experience
- 8+ years' experience in Software Engineering, Platform Engineering, DevOps, or MLOps.
- 5+ years' experience building and operating production MLOps platforms.
- Strong hands‑on experience with Azure-based MLOps architectures and AKS.
- Deep expertise in Kubernetes, containerisation, and model deployment patterns.
- Experience implementing monitoring, observability, and model lifecycle management.
- Hands‑on experience with CI/CD pipelines and Infrastructure as Code.
- Experience with Azure Monitor, Application Insights, Azure DevOps, and/or GitHub Actions.
- Proficiency with Terraform, Bicep, or equivalent.
- Strong Python and scripting skills.
- Experience supporting low-latency ML inference workloads and cloud cost optimisation initiatives.