Performance Engineer, GPU

Anthropic

New York (NY)

On-site

USD 280,000 - 850,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is seeking a GPU Performance Engineer to architect and optimize GPU systems powering Claude and related AI models. You will maximize GPU utilization, implement kernel-level optimizations, and scale pipelines across thousands of GPUs.

Ideal candidates will have deep GPU programming experience, a track record of performance improvements, and enjoy solving complex hardware-to-software challenges in a collaborative environment.

Qualifications

  • Experience with GPU programming and optimization at scale.
  • Ability to navigate complex hardware and ML frameworks.
  • Proven track record delivering measurable performance improvements.
  • Willingness to collaborate and pair program.

Responsibilities

  • Design and optimize GPU kernels for large models.
  • Develop distributed multi-node GPU training and inference pipelines.
  • Profile and optimize GPU utilization in production ML systems.
  • Collaborate with researchers and engineers to push AI infrastructure forward.

Skills

GPU programming
System optimization
Collaborative problem-solving
ML frameworks familiarity

Education

Bachelor’s degree

Tools

CUDA
Triton
CUTLASS
Flash Attention
Nsight

Job description

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role

Pioneering the next generation of AI requires breakthrough innovations in GPU performance and systems engineering. As a GPU Performance Engineer, you'll architect and implement the foundational systems that power Claude and push the frontiers of what's possible with large language models. You'll be responsible for maximizing GPU utilization and performance at unprecedented scale, developing cutting‑edge optimizations that directly enable new model capabilities and dramatically improve inference efficiency.

Working at the intersection of hardware and software, you'll implement state‑of‑the‑art techniques from custom kernel development to distributed system architectures. Your work will span the entire stack—from low‑level tensor core optimizations to orchestrating thousands of GPUs in perfect synchronization.

Strong candidates will have a track record of delivering transformative GPU performance improvements in production ML systems and will be excited to shape the future of AI infrastructure alongside world‑class researchers and engineers.

You might be a good fit if you
  • Have deep experience with GPU programming and optimization at scale
  • Are impact‑driven, passionate about delivering measurable performance breakthroughs
  • Can navigate complex systems from hardware interfaces to high‑level ML frameworks
  • Enjoy collaborative problem‑solving and pair programming
  • Want to work on state‑of‑the‑art language models with real‑world impact
  • Care about the societal impacts of your work
  • Thrive in ambiguous environments where you define the path forward
Strong candidates may also have experience with
  • GPU Kernel Development: CUDA, Triton, CUTLASS, Flash Attention, tensor core optimization
  • Performance Engineering: Kernel fusion, memory bandwidth optimization, profiling with Nsight
Representative projects
  • Co‑design attention mechanisms and algorithms for next‑generation hardware architectures
  • Develop custom kernels for emerging quantization formats and mixed‑precision techniques
  • Design distributed communication strategies for multi‑node GPU clusters
  • Optimize end‑to‑end training and inference pipelines for frontier language models
  • Build performance modeling frameworks to predict and optimize GPU utilization
  • Implement kernel fusion strategies to minimize memory bandwidth bottlenecks
  • Create resilient systems for planet‑scale distributed training infrastructure
  • Profile and eliminate performance bottlenecks in production serving infrastructure
  • Partner with hardware vendors to influence future accelerator capabilities and software stacks

The expected salary range for this position is $280,000 - $850,000 USD.

Logistics

Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience

Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience

Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Location‑based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Visa sponsorship: We do sponsor visas. If we make you an offer, we will make every reasonable effort to obtain a visa.

As set forth in Anthropic’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Performance Engineer, GPU
Performance Engineer, GPU

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive salary
Equity opportunities
Flexible working hours
+1
Performance Engineer, GPU
Performance Engineer, GPU

SignalAI • New York (NY)

Hybrid
USD 280,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+1
Performance Engineer
Performance Engineer

SignalAI • New York (NY)

Hybrid
USD 280,000 - 850,000
Optional equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Performance Engineer, Inference Systems San Francisco, CA | New York City, NY | Seattle, WA
Performance Engineer, Inference Systems San Francisco, CA | New York City, NY | Seattle, WA

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Engineering Manager, GPU (ML Accelerator)
Engineering Manager, GPU (ML Accelerator)

Anthropic • New York (NY)

On-site
USD 425,000 - 560,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
+2
Performance Engineer
Performance Engineer

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
Competitive compensation
Optional equity donation matching
Generous vacation and parental leave
+1
Engineering Manager, GPU (ML Accelerator)
Engineering Manager, GPU (ML Accelerator)

Anthropic • Seattle (WA)

On-site
USD 425,000 - 560,000
TPU Kernel Engineer
TPU Kernel Engineer

SignalAI • New York (NY)

Hybrid
USD 280,000 - 850,000
Staff+ Software Engineer, Inference Runtime
Staff+ Software Engineer, Inference Runtime

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 485,000
TPU Kernel Engineer Anthropic San Francisco, CA | New York City, NY | Seattle, WA
TPU Kernel Engineer Anthropic San Francisco, CA | New York City, NY | Seattle, WA

Neura Market • San Francisco (CA)

Hybrid
USD 280,000 - 850,000