Staff Cloud Native Engineer — AI GPU Infra Architect

Nscale

Greater London

On-site

GBP 110,000 - 170,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Nscale is hiring a Staff Cloud Native Software Engineer to build, operate, and improve the cloud-native software integrations that connect AI applications and networking components at scale.

In this role you’ll work on shared Kubernetes-based platforms, deployment patterns, observability foundations, infrastructure architecture, and operational tooling that help internal teams run services safely and efficiently on GPU-backed infrastructure.

Qualifications

  • At least 8 years of experience in production-level software development.
  • Deep hands-on experience building and operating Kubernetes-native software: custom controllers, operators, CRDs, or admission webhooks — using controller-runtime, client-go, or equivalent.
  • Strong understanding of Kubernetes internals: the API server, informer/lister patterns, reconciliation loops, and the object model.
  • Strong networking fundamentals — CNI, service mesh, kube-proxy/eBPF datapaths, DNS, load balancing — and experience building software that integrates with these systems.
  • Proficiency in Go (strongly preferred) or a similar language, with a track record of shipping well-tested, production-quality code at scale.
  • Experience with observability practices — metrics, tracing, structured logging — built into software rather than added afterward.
  • Comfortable owning components independently end-to-end, from design through operation, while setting direction for adjacent teams.
  • Experience with or strong interest in GPU-backed infrastructure and AI workload patterns is a plus.
  • Track record of leading technical design at a staff level and mentoring engineers through practical technical guidance.

Responsibilities

  • Design, build, and operate Kubernetes-native software — controllers, operators, CRDs, and admission webhooks — that connects AI applications with core networking components on GPU-backed infrastructure.
  • Extend Kubernetes control-plane capabilities to support AI workload requirements, including network policy controllers, CNI/service-mesh integrations, and resource/scheduling extensions.
  • Own significant components end-to-end and set the technical direction for how they're designed, deployed, and operated across the team.
  • Build reconciliation loops, informers, and client-go–based tooling that keep infrastructure state consistent between the API server, networking systems, and AI runtime components.
  • Develop operational tooling and automation that make Kubernetes-native services easier for internal teams to deploy, run, and support.
  • Drive infrastructure architecture decisions around how AI applications and networking components integrate across the platform, weighing trade-offs at a cross-team level.
  • Build observability foundations for controller and operator software — metrics, structured events, tracing, and status reporting surfaced through the Kubernetes API and platform dashboards.
  • Design systems that degrade gracefully and self-heal, using controller patterns (reconciliation, backoff, status conditions) to reduce manual intervention.
  • Debug and resolve complex issues spanning the Kubernetes control plane, networking (CNI, service mesh, kube-proxy/eBPF datapaths), and workload runtime behavior on GPU-backed infrastructure.
  • Define standards for safe rollout of controller and platform changes, including versioning, compatibility, and staged deployment.
  • Set technical direction for how the team builds Kubernetes-native software, establishing patterns for controller design, CRD schema evolution, and testing strategy.
  • Lead design discussions and code reviews, holding a high bar for Kubernetes API conventions and idiomatic client-go usage.
  • Partner with platform engineering, infrastructure, and product teams to translate real developer and operational needs into clean CRDs, APIs, and controller-managed abstractions.
  • Define reusable patterns, shared libraries, and scaffolding that let other teams build correctly on the platform without reinventing integration logic.
  • Mentor engineers in Kubernetes internals, controller-runtime patterns, and sound operational judgement.

Skills

Kubernetes expertise
Go programming
Software leadership
Code reviews

Education

Bachelor's degree in Computer Science or equivalent

Tools

controller-runtime
client-go
CRDs
admission webhooks

Job description

Nscale is hiring a Staff Cloud Native Software Engineer to build, operate, and improve the cloud-native software integrations that connect AI applications and networking components at scale.

In this role you’ll work on shared Kubernetes-based platforms, deployment patterns, observability foundations, infrastructure architecture, and operational tooling that help internal teams run services safely and efficiently on GPU-backed infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Cloud Native Software Engineer
Staff Cloud Native Software Engineer

Nscale • Greater London

On-site
GBP 110,000 - 170,000
Staff Engineer, Kubernetes GPU Inference Platform
Staff Engineer, Kubernetes GPU Inference Platform

Togetherai • Greater London

On-site
GBP 110,000 - 150,000
Senior Cloud & DevOps Architect — GPU-Accelerated AI/HPC
Senior Cloud & DevOps Architect — GPU-Accelerated AI/HPC

NVIDIA • United Kingdom

On-site
GBP 110,000 - 170,000
Staff HPC Systems Engineer Slurm & GPU Cloud Platform
Staff HPC Systems Engineer Slurm & GPU Cloud Platform

AI Startups UK • Greater London

Hybrid
GBP 110,000 - 170,000
Staff Software Engineer, Kubernetes-native GPU Inference
Staff Software Engineer, Kubernetes-native GPU Inference

Together AI • Greater London

Hybrid
GBP 100,000 - 160,000
Deployment Program Manager
Deployment Program Manager

Nscale • Greater London

On-site
GBP 100,000 - 130,000
Deployment Engineering Director, Systems Engineering
Deployment Engineering Director, Systems Engineering

Nscale • Greater London

On-site
GBP 120,000 - 180,000
Director of Customer Experience & AI Adoption
Director of Customer Experience & AI Adoption

Nscale • Trellech

On-site
GBP 126,000 - 189,000
Medical benefits
Dental benefits
Vision benefits
+3
Platform Engineer – Scale GPU Infra for AI Platform
Platform Engineer – Scale GPU Infra for AI Platform

Ineffable Intelligence LTD • Greater London

Hybrid
GBP 85,000 - 120,000
Director of Deployment Engineering - AI Cloud Platform
Director of Deployment Engineering - AI Cloud Platform

Nscale • United Kingdom

Remote
GBP 120,000 - 180,000