A complete application in a minute — tailored resume and cover letter, ready to send.
Anthropic seeks a GPU Performance Engineer to architect and implement foundational systems powering Claude, driving GPU utilization and optimization at unprecedented scale. You will work at the hardware-software boundary, from custom kernel development to distributed multi-node clusters.
Responsibilities include end-to-end optimization of training and inference pipelines, co-design of attention mechanisms for future architectures, and collaboration with hardware vendors to shape accelerator
Strong candidates will have a track record of delivering transformative GPU performance improvements in production ML systems and will be excited to shape the future of AI infrastructure alongside world-class researchers and engineersHave deep experience with GPU programming and optimization at scaleCare about the societal impacts of your workCan navigate complex systems from hardware interfaces to high-level ML frameworksAre impact-driven, passionate about delivering measurable performance breakthroughsEnjoy collaborative problem-solving and pair programmingThrive in ambiguous environments where you define the path forwardWant to work on state-of-the-art language models with real-world impactEducation requirements: We require at least a Bachelor's degree in a related field or equivalent experienceGPU Kernel Development: CUDA, Triton, CUTLASS, Flash Attention, tensor core optimizationML Compilers & Frameworks: PyTorch/JAX internals, torch.compile, XLA, custom operatorsPerformance Engineering: Kernel fusion, memory bandwidth optimization, profiling with NsightDistributed Systems: NCCL, NVLink, collective communication, model parallelismLow-Precision: INT8/FP8 quantization, mixed-precision techniquesProduction Systems: Large-scale training infrastructure, fault tolerance, cluster orchestration