Performance Engineer (GPU)

Anthropic

York and North Yorkshire

On-site

GBP 90,000 - 140,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Comprehensive health insurance
Fertility benefits
22 weeks parental leave
Flexible paid time off

Job summary

Anthropic seeks a GPU Performance Engineer to architect and implement foundational systems powering Claude, driving GPU utilization and optimization at unprecedented scale. You will work at the hardware-software boundary, from custom kernel development to distributed multi-node clusters.

Responsibilities include end-to-end optimization of training and inference pipelines, co-design of attention mechanisms for future architectures, and collaboration with hardware vendors to shape accelerator

Qualifications

  • Track record delivering GPU performance improvements in production ML systems.
  • Experience optimizing end-to-end training and inference pipelines at scale.
  • Strong understanding of modern ML frameworks and low-level GPU hardware.

Responsibilities

  • Architect and implement foundational GPU performance systems for Claude.
  • Maximize GPU utilization and inference efficiency at scale.
  • Develop kernel-level optimizations and distributed architectures across thousands of GPUs.
  • Co-design attention mechanisms for next-gen hardware architectures.
  • Collaborate with hardware vendors to influence accelerator capabilities.

Skills

GPU performance optimization
Collaborative problem-solving
Hardware-software co-design

Education

Bachelor's degree or equivalent experience

Tools

CUDA
Triton
CUTLASS
Flash Attention
Tensor cores
PyTorch/JAX internals
torch.compile
XLA
NCCL
NVLink

Job description

  • As a GPU Performance Engineer, you'll architect and implement the foundational systems that power Claude and push the frontiers of what's possible with large language models
  • You'll be responsible for maximizing GPU utilization and performance at unprecedented scale, developing cutting-edge optimizations that directly enable new model capabilities and dramatically improve inference efficiency
  • Working at the intersection of hardware and software, you'll implement state-of-the-art techniques from custom kernel development to distributed system architectures
  • Your work will span the entire stack—from low-level tensor core optimizations to orchestrating thousands of GPUs in perfect synchronization
  • Co-design attention mechanisms and algorithms for next-generation hardware architectures
  • Develop custom kernels for emerging quantization formats and mixed-precision techniques
  • Design distributed communication strategies for multi-node GPU clusters
  • Optimize end-to-end training and inference pipelines for frontier language models
  • Build performance modeling frameworks to predict and optimize GPU utilization
  • Implement kernel fusion strategies to minimize memory bandwidth bottlenecks
  • Create resilient systems for planet-scale distributed training infrastructure
  • Profile and eliminate performance bottlenecks in production serving infrastructure
  • Partner with hardware vendors to influence future accelerator capabilities and software stacks
Benefits
  • Comprehensive health, dental, and vision insurance for you and your dependents
  • Inclusive fertility benefits via Carrot Fertility
  • 22 weeks of paid parental leave
  • Flexible paid time off and absence policies
  • Mental health support for you and your dependents
  • Competitive salary and equity packages
  • Optional equity donation matching at a 1:1 ratio, up to 25% of your equity grant
  • Retirement plans with competitive matching
  • Life and income protection plans
  • $500/month flexible wellness and time saver stipend
  • Commuter benefits
  • Annual education stipend
  • Home office stipends
  • Relocation support for those moving for Anthropic
  • Daily meals and snacks in the office

Strong candidates will have a track record of delivering transformative GPU performance improvements in production ML systems and will be excited to shape the future of AI infrastructure alongside world-class researchers and engineersHave deep experience with GPU programming and optimization at scaleCare about the societal impacts of your workCan navigate complex systems from hardware interfaces to high-level ML frameworksAre impact-driven, passionate about delivering measurable performance breakthroughsEnjoy collaborative problem-solving and pair programmingThrive in ambiguous environments where you define the path forwardWant to work on state-of-the-art language models with real-world impactEducation requirements: We require at least a Bachelor's degree in a related field or equivalent experienceGPU Kernel Development: CUDA, Triton, CUTLASS, Flash Attention, tensor core optimizationML Compilers & Frameworks: PyTorch/JAX internals, torch.compile, XLA, custom operatorsPerformance Engineering: Kernel fusion, memory bandwidth optimization, profiling with NsightDistributed Systems: NCCL, NVLink, collective communication, model parallelismLow-Precision: INT8/FP8 quantization, mixed-precision techniquesProduction Systems: Large-scale training infrastructure, fault tolerance, cluster orchestration

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Performance Engineer
Performance Engineer

Anthropic • York and North Yorkshire

On-site
GBP 110,000 - 150,000
Health insurance
Fertility benefits
Parental leave 22 weeks
+12
Engineering Manager (GPU, ML Accelerator)
Engineering Manager (GPU, ML Accelerator)

Anthropic • York and North Yorkshire

On-site
GBP 110,000 - 150,000
Health insurance
Dental insurance
Vision insurance
+15
TPU Kernel Engineer
TPU Kernel Engineer

Anthropic • York and North Yorkshire

On-site
GBP 110,000 - 150,000
Health insurance
Dental & vision
Parental leave 22 weeks
+7
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Member of Technical Staff - Mid-Training Infra
Member of Technical Staff - Mid-Training Infra

Reflection • Greater London

On-site
GBP 70,000 - 100,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
Training / AI Infrastructure Engineering & Research London
Training / AI Infrastructure Engineering & Research London

Genesis • Greater London

Hybrid
GBP 90,000 - 130,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Quant Blueprint LLC • Greater London

On-site
GBP 50,000 - 70,000
Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid)
Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid)

Pure Resourcing Solutions Limited • Linton

Hybrid
GBP 90,000 - 120,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

CVFine by Instrovate Technologies • Greater London

On-site
GBP 70,000 - 90,000
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1