DevOps and Machine Learning Operations Engineer

Kintec Global Recruitment

Manchester

On-site

GBP 70,000 - 90,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Kintec Global Recruitment seeks a DevOps and Machine Learning Operations Engineer to join a specialist software team in Manchester. You will build and operate delivery pipelines, environments, and the model-serving path with IaC, CI/CD, observability, cost control, and production support across the platform.

You will bring production experience, strong Terraform and Kubernetes skills, CI/CD ownership, and the ability to discuss cost and architecture with engineers and business stakeholders.

Qualifications

  • Proven production experience in DevOps or platform engineering with reliability accountability.
  • Strong infrastructure as code practice (Terraform) and container orchestration (Kubernetes).
  • Demonstrated CI/CD ownership across build, test, security gates, deployment, and rollback.
  • Observability experience: structured logging, metrics, tracing, alerting, and on-call readiness.
  • Experience deploying and operating ML models in production, including versioning and monitoring.
  • Proficiency in Python, Go, or TypeScript for automation and tooling.

Responsibilities

  • Define and maintain cloud infrastructure as code for development, staging, and production.
  • Build and manage delivery pipelines with automated testing, security gates, and rollback.
  • Instrument services with logging, metrics, and tracing; define and monitor SLAs.
  • Run production support and incident response with post-incident reviews.
  • Package, version, and deploy ML models; monitor drift, latency, and cost in serving.
  • Address cost and architectural trade-offs with stakeholders.

Job description

DevOps and Machine Learning Operations Engineer

Location: Manchester, UK
Contract Type: Contract Position

About the Project

Join a specialist software development team delivering a new, business-critical technology platform for an established international organisation. This greenfield development project involves modern cloud architecture, data-intensive applications, and AI-enabled capabilities. You will work as part of a multidisciplinary team alongside experienced software, AI, infrastructure, and design professionals, with direct involvement in taking the platform from development through to production. Security, scalability, maintainability, data protection, and production readiness are key priorities throughout the project.

Role Purpose

As a DevOps and MLOps Engineer, you will build and operate the delivery pipelines, environments, and runtime platform the project depends on, as well as the model-serving path. The role covers infrastructure as code, continuous integration and deployment, observability, cost control, and production support. Applicants must have experience running systems in production and being accountable for their reliability.

Main Responsibilities
  • Infrastructure and Environments:
    • Define and maintain cloud infrastructure as code, ensuring reviewable changes and reproducible environments for development, staging, and production.
    • Manage secrets, certificates, network boundaries, and access with least privilege as the default and auditable grants.
    • Maintain environment consistency and make any necessary differences explicit.
  • Continuous Integration and Deployment:
    • Build pipelines that test, scan, build, and deploy, with quality gates that fail closed.
    • Automate database migration and rollback for reversible releases.
    • Support progressive delivery, including staged rollout and fast rollback, with deployment tracking.
  • Observability and Operations:
    • Instrument services with structured logging, metrics, and tracing, and define user-focused alerts.
    • Establish service level objectives and report against them.
    • Run incident response: triage, mitigate, restore, and document post-incident reviews with corrective actions.
  • Machine Learning Operations:
    • Package, version, and deploy models and dependencies for traceable predictions.
    • Automate evaluation before promotion and monitor drift, latency, and cost during serving.
    • Enable safe rollback of models independently of application releases.
    • Manage inference cost and capacity, including batching, caching, and hardware selection.
  • Security and Compliance:
    • Apply dependency and container scanning, patching, and image provenance as part of the pipeline.
    • Support data protection obligations, including retention, deletion, and access logging.
    • Document and rehearse recovery objectives.
Essential Technical Experience
  • Substantial professional experience in DevOps, platform, or site reliability engineering on production systems with accountability for reliability.
  • Strong infrastructure as code practice (e.g., Terraform) and experience with a container orchestration platform such as Kubernetes.
  • Demonstrated CI/CD ownership, including build, test, security gates, deployment, and rollback.
  • Practical observability experience: structured logging, metrics, tracing, alert design, and on-call.
  • Experience deploying and operating machine learning models in production, including versioning, evaluation, and monitoring.
  • Competence in at least one of Python, Go, or TypeScript for automation and tooling.
  • Sound understanding of cloud networking, identity and access management, and secret handling.
  • Ability to reason about cost and explain architectural trade-offs to both engineers and business stakeholders.
Desirable Experience
  • Experience with GPU scheduling and inference optimisation.
  • Experience with feature stores, vector databases, or retrieval pipelines.
  • Exposure to regulated environments and formal audit.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Ops Engineer 201043
Senior ML Ops Engineer 201043

Harnham • United Kingdom

On-site
GBP 70,000 - 110,000
Ongoing training
Ownership of ML Ops function
DevOps Engineer
DevOps Engineer

SoCode Recruitment • Cambridge

On-site
GBP 50,000 - 70,000
MLOps Engineer
MLOps Engineer

DGH Recruitment • City Of London

On-site
GBP 42,000 - 78,000
Senior ML Engineer
Senior ML Engineer

Artificial Intelligence Jobs • Greater London

Hybrid
GBP 75,000 - 85,000
Annual bonus (10%)
MLOps Engineer
MLOps Engineer

DGH Recruitment Ltd • City Of London

On-site
GBP 70,000 - 110,000
Senior ML Ops Engineer
Senior ML Ops Engineer

Harnham - Data and Analytics Recruitment • Greater London

Hybrid
GBP 75,000 - 85,000
Hybrid work model
MLOps Engineer
MLOps Engineer

Harnham - Data & Analytics Recruitment • Greater London

Hybrid
GBP 75,000 - 85,000
Lead MLOps Engineer
Lead MLOps Engineer

Gravitas Recruitment Group (Global) Ltd • Manchester

On-site
GBP 80,000 - 100,000
Senior ML Ops Engineer
Senior ML Ops Engineer

Harnham • Greater London

Hybrid
GBP 75,000 - 85,000
Private healthcare
DevOps Engineer
DevOps Engineer

Matchtech • Lancashire

Hybrid
GBP 96,000 - 129,000