Senior MLOps Engineer
London (Hybrid) | Salary: up to £85,000
This is an opportunity to join a growing AI and data organisation where your work will directly support the deployment, scalability, and reliability of innovative machine learning solutions. You will play a pivotal role in building the infrastructure that enables advanced AI services to deliver real-world impact at scale.
The Company
They are a mission-led technology organisation that uses data, machine learning, and AI to help organisations make better decisions and deliver meaningful outcomes. With a growing client base and continued investment in their platform, they are expanding their engineering capabilities to support the next stage of growth. Their environment combines technical excellence with a strong focus on responsible AI, innovation, and continuous improvement. You will be joining a collaborative team where your expertise will help shape both the platform and the wider AI strategy.
The Role
- Build, maintain, and evolve the cloud infrastructure that supports machine learning and AI services in production.
- Deploy, manage, and monitor machine learning models using Azure ML, Azure Kubernetes Service (AKS), and related Azure services.
- Design and develop robust orchestration pipelines for model training, deployment, inference, and retraining workflows.
- Implement infrastructure-as-code solutions using tools such as Terraform, Bicep, or similar technologies.
- Establish monitoring, logging, alerting, and observability frameworks to ensure reliability and performance.
- Support scalable deployment across multiple client environments while maintaining security and operational excellence.
- Collaborate with data scientists, engineers, and wider business stakeholders to enable successful AI delivery.
- Contribute to AI governance, best practices, and responsible AI principles across the organisation.
Your Skills & Experience
- Strong commercial experience within MLOps, Platform Engineering, or Infrastructure Engineering supporting machine learning systems.
- Advanced Python programming skills with a strong software engineering mindset.
- Hands‑on experience with Azure-native MLOps services, including model deployment, pipelines, environments, and compute resources.
- Proven expertise deploying and managing containerised applications on Kubernetes.
- Experience building and maintaining CI/CD pipelines using Azure DevOps or similar tools.
- Knowledge of infrastructure-as-code approaches using Terraform, Bicep, Pulumi, or equivalent technologies.
- Experience implementing monitoring and observability solutions using tools such as Prometheus, Grafana, or similar.
- Familiarity with orchestration platforms including Dagster, Airflow, Prefect, or related technologies.
Desirable experience includes:
- Model serving infrastructure for real-time or batch inference workloads.
- Multi-tenant platform environments.
- Vector databases such as Qdrant.
- Development of high-performance APIs and data-intensive services.
- Broader Azure ecosystem expertise including Azure Container Registry, Azure Blob Storage, Azure Monitor, and Azure Key Vault.
What They Offer
- The opportunity to work on innovative AI and machine learning projects with genuine societal impact.