TPU Kernel Engineer for High-Performance ML Systems

SignalAI

New York (NY)

Hybrid

USD 280,000 - 850,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is seeking a TPU Kernel Engineer to identify and address performance issues across ML systems, including research, training, and inference. You will design and optimize kernels for the TPU and provide feedback to researchers about how model changes impact performance.

Strong candidates will have a track record of solving large-scale systems problems and low-level optimization, with experience on TPUs, GPUs, or accelerators, and an interest in ML research and societal impact.

Qualifications

  • Experience optimizing ML systems for TPUs, GPUs, or accelerators.
  • Bias towards flexibility and impact in work.
  • Willing to take on tasks outside job description.
  • Enjoy pair programming and collaboration.
  • Interest in machine learning research.
  • Concern for societal impacts of AI work.

Responsibilities

  • Implement low-latency, high-throughput sampling for large language models.
  • Adapt existing models for low-precision inference.
  • Build quantitative models of system performance.
  • Design and implement custom collective communication algorithms.
  • Debug kernel performance at the assembly level.

Skills

TPU optimization
Results oriented
Pair programming
ML research interest
Societal impact awareness

Education

Bachelor's degree

Tools

Low-level kernel optimization
ML framework internals
Transformer architectures
Computer architecture

Job description

Anthropic is seeking a TPU Kernel Engineer to identify and address performance issues across ML systems, including research, training, and inference. You will design and optimize kernels for the TPU and provide feedback to researchers about how model changes impact performance.

Strong candidates will have a track record of solving large-scale systems problems and low-level optimization, with experience on TPUs, GPUs, or accelerators, and an interest in ML research and societal impact.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid TPU Kernel Engineer for High-Performance ML
Hybrid TPU Kernel Engineer for High-Performance ML

Neura Market • San Francisco (CA)

Hybrid
USD 280,000 - 850,000
TPU Kernel Engineer — Lead Low-Latency ML Kernels (Hybrid)
TPU Kernel Engineer — Lead Low-Latency ML Kernels (Hybrid)

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
TPU Kernel Engineer
TPU Kernel Engineer

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
TPU Kernel Architect
TPU Kernel Architect

Anthropic • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Engineering Manager: ML Performance & TPU Optimizations
Engineering Manager: ML Performance & TPU Optimizations

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Senior TPU Performance Architect for AI Systems
Senior TPU Performance Architect for AI Systems

Socket.dev • Sunnyvale (CA)

Hybrid
USD 174,000 - 252,000
TPU Systems Engineer — High-Performance ML Inference Equity
TPU Systems Engineer — High-Performance ML Inference Equity

RadixArk • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Significant founding team equity
Comprehensive health benefits
Flexible work arrangements
Lead Kernel Engineer/Architect (m/f/d)
Lead Kernel Engineer/Architect (m/f/d)

EPAM Systems • Germany (OH)

Hybrid
USD 104,000 - 152,000
Senior ML Compute Efficiency Engineer (GPU/TPU Performance)
Senior ML Compute Efficiency Engineer (GPU/TPU Performance)

Socket.dev • Santa Clara (CA)

On-site
USD 130,000 - 180,000
GPU Kernel Engineer — Fast ML Training & Inference
GPU Kernel Engineer — Fast ML Training & Inference

Tilde Research • Palo Alto (CA)

On-site
USD 150,000 - 260,000