High-Throughput ML Systems Performance Engineer

Anthropic

York and North Yorkshire

On-site

GBP 110,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
Fertility benefits
Parental leave 22 weeks
Flexible PTO
Mental health support
Equity package
Equity donation matching
Retirement plans
Life and income protection
Wellness stipend
Commuter benefits
Education stipend
Home office stipend
Relocation support
In-office meals

Job summary

Anthropic in the United Kingdom is seeking a Performance Engineer to identify novel systems problems and build high-throughput, low-latency distributed services for large language models. You will implement sampling, GPU kernel adaptations for low-precision inference, custom load-balancing, and fault-tolerant architectures, collaborating with teams to maximize throughput and reliability.

This role requires deep systems expertise at supercomputing scale and a readiness to pair-program and own

Qualifications

  • Significant software engineering or ML experience, especially at scale.
  • Experience solving large-scale systems problems.
  • Interest in ML and societal impact, with results-driven mindset.

Responsibilities

  • Implement low-latency high-throughput sampling for large language models.
  • Implement GPU kernels to adapt models to low-precision inference.
  • Write a custom load-balancing algorithm to optimize serving efficiency.
  • Build quantitative models of system performance.
  • Design and implement fault-tolerant distributed systems with complex networks.
  • Debug kernel-level network latency spikes in containerized environments.

Skills

Low-latency systems
GPU kernel programming
Distributed systems
Load balancing
Kernel debugging
Performance modeling

Tools

Docker
Kubernetes

Job description

Anthropic in the United Kingdom is seeking a Performance Engineer to identify novel systems problems and build high-throughput, low-latency distributed services for large language models. You will implement sampling, GPU kernel adaptations for low-precision inference, custom load-balancing, and fault-tolerant architectures, collaborating with teams to maximize throughput and reliability.

This role requires deep systems expertise at supercomputing scale and a readiness to pair-program and own

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

TPU Kernel Engineer: Optimize ML Systems at Scale
TPU Kernel Engineer: Optimize ML Systems at Scale

Anthropic • York and North Yorkshire

On-site
GBP 110,000 - 150,000
Health insurance
Dental & vision
Parental leave 22 weeks
+7
Performance Engineer
Performance Engineer

Anthropic • York and North Yorkshire

On-site
GBP 110,000 - 150,000
Health insurance
Fertility benefits
Parental leave 22 weeks
+12
Real-Time ML Systems Performance Engineer
Real-Time ML Systems Performance Engineer

Janestreet • Greater London

On-site
GBP 70,000 - 90,000
ML Systems Performance Engineer
ML Systems Performance Engineer

Quant Blueprint LLC • Greater London

On-site
GBP 50,000 - 70,000
Engineering Manager, GPU & ML Systems Scaling
Engineering Manager, GPU & ML Systems Scaling

Anthropic • York and North Yorkshire

On-site
GBP 110,000 - 150,000
Health insurance
Dental insurance
Vision insurance
+15
Remote Performance Engineer: ML Training & Kernels
Remote Performance Engineer: ML Training & Kernels

Cohere • Greater London

On-site
GBP 75,000 - 95,000
Co-working benefit
Daily lunch program
Regular community and social events
Senior ML Systems Engineer: Large-Scale Performance
Senior ML Systems Engineer: Large-Scale Performance

Stanford Black Limited • England

On-site
GBP 70,000 - 150,000
LLM Performance Engineer — Scale, Optimize & Deploy
LLM Performance Engineer — Scale, Optimize & Deploy

Isomorphic Labs • Greater London

Hybrid
GBP 120,000 - 160,000
ML Performance Engineer – Scale GPU/CPU Workloads
ML Performance Engineer – Scale GPU/CPU Workloads

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Lunch provided
35 days’ annual leave
9% company pension contributions
+4
Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1