Senior AI Cloud Platform Engineer

Nscale

Greater London

On-site

GBP 60,000 - 90,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Nscale is seeking a Platform Operations & Engineering professional to own and evolve a Kubernetes-based infrastructure for AI workloads. You’ll drive deployment workflows, observability, and automation across production platforms, collaborating with software, infra, and SRE teams to deliver reliable cloud-native capabilities.

You will improve incident prevention, rollout safety, and platform resilience while mentoring junior engineers and shaping best practices for service deployment and runtime

Qualifications

  • Strong hands-on experience operating and improving Kubernetes-based platforms in production.
  • Solid experience with infrastructure automation, CI/CD, configuration management, or GitOps-style workflows.
  • Strong understanding of reliability engineering principles, including observability, incident response, failure analysis, and operational readiness.
  • Experience writing production-quality automation, tooling, or backend code in Go, Python, Bash, or similar languages.
  • Good Linux fundamentals, including processes, filesystems, cgroups, service behaviour, and system debugging.
  • Good networking fundamentals, including TCP/IP, DNS, routing, load balancing, and container or overlay networking concepts.
  • Experience debugging complex production issues across multiple system layers.
  • Ability to work independently on substantial technical problems while collaborating effectively with adjacent teams.
  • Experience mentoring or supporting less experienced engineers through practical technical guidance.

Responsibilities

  • Build and improve shared cloud-native platform capabilities used by internal engineering teams to run AI applications and services.
  • Own significant parts of the platform area, including Kubernetes cluster operations, workload runtime configuration, deployment workflows, observability foundations, or environment automation.
  • Improve the reliability, scalability, and supportability of platform services through practical engineering and operational enhancements.
  • Develop automation, tooling, and configuration that reduce manual effort, improve consistency, and make the platform easier to use and operate.
  • Apply software engineering where it creates leverage, including scripts, services, CI/CD automation, operational tooling, and platform integrations.
  • Improve incident prevention, detection, response, and recovery across the platform areas you support.
  • Build and refine observability for platform services, including metrics, logs, tracing, dashboards, alerts, and other useful operational signals.
  • Strengthen rollout safety, capacity awareness, failure handling, and recovery procedures for production environments.
  • Debug and resolve complex issues spanning Kubernetes, Linux, networking, storage, workload runtime behaviour, and cloud or datacentre infrastructure dependencies.
  • Enhance operational playbooks, runbooks, and engineering practices to reduce toil and increase service resilience.
  • Contribute to design discussions, code reviews, and operational standards within the platform engineering team.
  • Collaborate with software engineering, infrastructure, and SRE teams to deliver platform capabilities that are practical, supportable, and aligned to operational needs.
  • Define sensible defaults, paved roads, and supportable patterns for service deployment and runtime operations.
  • Mentor less experienced engineers in platform engineering fundamentals, operational judgement, and good automation practices.

Skills

Kubernetes
CI/CD
Go
Python
Bash
Linux
Networking
GitOps
Observability
Incident response
Automation
SRE

Tools

Terraform

Job description

Nscale is seeking a Platform Operations & Engineering professional to own and evolve a Kubernetes-based infrastructure for AI workloads. You’ll drive deployment workflows, observability, and automation across production platforms, collaborating with software, infra, and SRE teams to deliver reliable cloud-native capabilities.

You will improve incident prevention, rollout safety, and platform resilience while mentoring junior engineers and shaping best practices for service deployment and runtime

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Platform Engineer, Kubernetes & Automation
Senior Cloud Platform Engineer, Kubernetes & Automation

nscaleoperationsukltd • United Kingdom

On-site
GBP 70,000 - 110,000
Staff Cloud-Native Engineer: Kubernetes & AI Infra Leader
Staff Cloud-Native Engineer: Kubernetes & AI Infra Leader

nscaleoperationsukltd • United Kingdom

On-site
GBP 90,000 - 150,000
Director of Deployment Engineering - AI Cloud Platform
Director of Deployment Engineering - AI Cloud Platform

Nscale • United Kingdom

Remote
GBP 120,000 - 180,000
Senior Cloud Native Platform Engineer
Senior Cloud Native Platform Engineer

Nscale • Greater London

On-site
GBP 60,000 - 90,000
Senior DevOps Engineer - Cloud, Kubernetes & AI Platform
Senior DevOps Engineer - Cloud, Kubernetes & AI Platform

SBS • Greater London

On-site
GBP 85,000 - 120,000
Health coverage
Retirement plans
Paid time off
+4
Lead AI Infrastructure & Platform Engineering
Lead AI Infrastructure & Platform Engineering

Scale AI, Inc. • Greater London

On-site
GBP 110,000 - 160,000
Kubernetes Infrastructure Engineer for AI Software
Kubernetes Infrastructure Engineer for AI Software

EngineersOfAI • Greater London

On-site
GBP 50,000 - 70,000
Kubernetes Infrastructure Engineer for AI Software
Kubernetes Infrastructure Engineer for AI Software

EngineersOfAI • Cambridge

On-site
GBP 45,000 - 65,000
Senior AI Platform Architect – Hybrid
Senior AI Platform Architect – Hybrid

Webhosting • England

Hybrid
GBP 110,000 - 140,000
Senior AI Platform Engineer - Hybrid, Scalable & Secure
Senior AI Platform Engineer - Hybrid, Scalable & Secure

Camwebdir • United Kingdom

Hybrid
GBP 90,000 - 140,000