Staff Software Engineer, Kubernetes Platform

Humanloop

Greater London

Hybrid

GBP 57,000 - 73,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity donation matching
Generous vacation
Parental leave
Flexible working hours
Collaborative office space

Job summary

Anthropic seeks a senior systems/engineering leader to own and extend the Kubernetes scheduler for accelerator fleets, building custom plugins and scaling the control plane. You’ll design core cluster services and collaborate with research, training, and inference teams to translate workload needs into platform capabilities.

You will work with cloud providers, participate in on-call duties, and help establish postmortems, runbooks, and SLOs to prevent repeated failures.

Qualifications

  • Significant software engineering experience building and operating production distributed systems.
  • Proficiency in Go, Python, Rust, or C++.
  • Deep Kubernetes experience including schedulers, controllers, apiserver, or large multi-tenant clusters.
  • Strong debugging across stack from API to network-level root causes.
  • Design for reliability, correctness and clear failure semantics.

Responsibilities

  • Own, operate, and extend the Kubernetes scheduler for accelerator fleets with custom scheduling plugins.
  • Scale the Kubernetes control plane to support large clusters.
  • Design, build, and operate core cluster services used by workloads.
  • Build and maintain custom controllers, operators, and CRDs.
  • Collaborate with research, training, and inference to translate workload requirements into platform capabilities.
  • Partner with cloud providers on required features and escalations.
  • Participate in on-call and incident response; design postmortems, runbooks, and SLOs.

Skills

Distributed systems
Go
Python
Rust
C++
Debugging complex issues
Reliability design
Communication
Leadership
Bachelors degree

Education

Bachelors degree or equivalent

Tools

Kubernetes
etcd
ZooKeeper
Consul
kube-scheduler
apiserver
controller-runtime

Job description

Salary: £57,000 - 73,000 per year
Requirements
  • Significant software engineering experience building and operating production distributed systems
  • Proficiency in at least one systems-appropriate language such as Go, Python, Rust, or C++
  • Deep, hands-on Kubernetes experience beyond basic usage, including scheduler, controllers, apiserver, or operating large multi-tenant clusters
  • Demonstrated ability to debug complex issues across the stack, from API behavior to node- and network-level root causes
  • A track record of designing for reliability, correctness, and clear failure semantics in systems other engineers depend on
  • Strong written and verbal communication skills, with comfort building consensus with internal stakeholders
  • Experience with Kubernetes internals or contributions such as kube-scheduler, the scheduling framework, apiserver, etcd, client-go, controller-runtime, or similar
  • Experience building or operating cluster schedulers or batch systems such as Kueue, Volcano, Slurm, or in-house equivalents
  • Background scaling control planes or coordination systems such as etcd, ZooKeeper, Consul, or large DNS/service-mesh deployments
  • Familiarity with ML infrastructure such as GPUs, TPUs, or Trainium; gang scheduling; topology-aware placement; or collective networking such as NCCL
  • Experience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as Code
  • Low-level systems experience such as Linux kernel tuning, cgroups, or eBPF
  • 12+ years of relevant industry experience, including time leading large, ambiguous infrastructure projects
  • Bachelors degree or an equivalent combination of education, training, and/or experience
  • A field relevant to the role as demonstrated through coursework, training, or professional experience
Responsibilities
  • Own, operate, and extend the Kubernetes scheduler for our accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption
  • Scale the Kubernetes control plane, including apiserver, etcd, and controller-manager, to support clusters far beyond typical limits, and identify the next bottleneck before it finds us
  • Design, build, and operate core cluster services such as service discovery that every workload in the fleet depends on
  • Build and maintain custom controllers, operators, and CRDs
  • Partner with research, training, and inference to understand workload shapes and translate requirements into platform capabilities
  • Collaborate with cloud providers on required features and escalations
  • Participate in on-call, lead incident response, and design processes such as postmortems, runbooks, and SLOs to help the team avoid repeating failures
Technologies
  • AI
  • API
  • AWS
  • Cloud
  • GCP
  • Support
  • Kubernetes
  • Linux
  • Network
  • Python
  • Rust
  • ZooKeeper
  • NodeJS
More

We are Anthropic, a public benefit corporation headquartered in San Francisco, building reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. Our Kubernetes Platform team runs one of the industrys largest AI compute fleets across multiple cloud providers and datacenters, and we own the control plane that keeps it operating at scale. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a collaborative office space. We also have a location-based hybrid policy requiring staff to be in one of our offices at least 25% of the time, and we sponsor visas on a case-by-case basis where possible.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Kubernetes Platform Engineer | Hybrid & Flexible Hours
Staff Kubernetes Platform Engineer | Hybrid & Flexible Hours

Humanloop • Greater London

Hybrid
GBP 57,000 - 73,000
Equity donation matching
Generous vacation
Parental leave
+2
Staff Software Engineer, Infrastructure (Distributed Systems)
Staff Software Engineer, Infrastructure (Distributed Systems)

Anthropic Limited • Greater London

Hybrid
GBP 110,000 - 170,000
Kubernetes Platform Engineer
Kubernetes Platform Engineer

Barlowe LLP • Greater London

On-site
GBP 65,000 - 85,000
Highly competitive compensation plus annual discretionary bonus
Lunch provided via Just Eat for Business
30 days’ annual leave
+4
Senior Engineering Manager (Capacity Engineering)
Senior Engineering Manager (Capacity Engineering)

Anthropic • York and North Yorkshire

On-site
GBP 150,000 - 190,000
Health, dental, and vision insurance
Fertility benefits via Carrot Fertilty
Parental leave (22 weeks)
Kubernetes Platform Engineer
Kubernetes Platform Engineer

G-Research • Greater London

On-site
GBP 90,000 - 120,000
Highly competitive compensation
Annual discretionary bonus
Lunch provided
+5
Software Engineer (Kubernetes)
Software Engineer (Kubernetes)

CoreWeave • York and North Yorkshire

On-site
GBP 90,000 - 150,000
Staff Software Engineer, Inference / Compute Infrastructure Engineering
Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together AI • Greater London

Hybrid
GBP 100,000 - 160,000
Staff Software Engineer, Infrastructure (Distributed Systems)
Staff Software Engineer, Infrastructure (Distributed Systems)

AI Startups UK • Greater London

Hybrid
GBP 325,000 - 390,000
Competitive compensation
Equity donation matching (optional)
Generous vacation and parental leave
+1
Staff Software Engineer, Observability & Profiling
Staff Software Engineer, Observability & Profiling

AI Chopping Block • Greater London

Hybrid
GBP 325,000 - 390,000
Competitive compensation
Equity donation matching
Generous vacation
+3
Staff Software Engineer, Observability & Profiling
Staff Software Engineer, Observability & Profiling

AI Startups UK • Greater London

Hybrid
GBP 325,000 - 390,000
Office space
Flexible working hours
Generous vacation
+2