Performance Engineer, Inference Engine

EngineersOfAI

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Anthropic is seeking a Performance Engineer for its Inference Engine to optimize LLM processing across accelerator and cloud platforms. You will work on throughput, latency, and cost, ensuring robust, scalable performance from host to device and across chips.

Experience with Rust/C++ and GPU/accelerator programming is essential. The role emphasizes strong systems thinking, observability, and collaboration in a fast-moving AI safety-focused environment.

Qualifications

  • LLM inference: prefill and decode on accelerator compute.
  • Proven quick learner: ramped fast in deep, unfamiliar systems.
  • Strong systems programming (Rust, C++, or similar).
  • Analytical about performance: observe, hypothesize, test, then measure.
  • Low ego: receptive to feedback and collaboration.
  • Enjoy pair programming and consider societal impact of work.

Responsibilities

  • Keep device utilization high; accelerators should not wait due to overhead.
  • Reuse state to avoid recompute when cheaper than recomputing.
  • Build observability to model impact of potential improvements.
  • Deploy improvements across hosts and devices while maintaining safety.

Skills

LLM inference
Rust
C++
systems programming
performance analysis
pair programming
analytical mindset
memory optimization
GPU/Accelerator programming

Tools

CUDA
RDMA
PCIe

Job description

About Anthropic

Anthropics mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

Performance Engineer, Inference Engine
About the Role

Anthropics inference engine is the software between the accelerator kernels and the routing layer. It manages the entire token path in between: batching requests, laying the model out across chips, managing memory for weights and activations, coordinating every forward pass, and managing model state across requests. Built in-house, it runs on all of our accelerator platforms, serving Claude to millions of users and running our research workloads.

You will work on building and optimizing this system at Anthropics scale: improving throughput, cost, reliability, and latency across all accelerator and cloud platforms. You are intimately familiar with the hardware and bandwidth numbers (FLOPs, HBM, PCIe, RDMA, network links, etc.) and can model a problem quickly: where the time and bytes go, and what sets the bound. The role is deeply technical and high-impact, and suits engineers who enjoy working across accelerator programming, high-performance systems that seamlessly coordinate between host and device, and large-scale distributed systems. Familiarity with the transformer architecture is a plus.

Some example recurring themes:

  • Keep device utilization high. Accelerators should never be waiting due to other overheads.

  • Reuse instead of recompute. Keep model state cached and reuse it whenever that is cheaper than computing it again.

  • Measure, model, then change. We build the observability to see where the gaps are, model the impact of potential improvements, deploy them, and go around again, with Claude speeding up every turn of that loop.

  • Tokens you can trust. Ensuring model quality matters more than efficiency. We build the infrastructure to ensure Claude maintains its intelligence across platforms and over time.

  • Safety on every token. We work closely with our safeguards and safety teams. The inference engine is the backbone behind our production safety systems, ensuring efficiency without compromising robustness.

Minimum Qualifications
  • A working mental model of LLM inference: how prefill and decode land on an accelerator's compute, memory, and interconnect, and what the host is doing meanwhile

  • Proven quick learner: ramped fast in deep, unfamiliar systems and shipped consequential changes quickly

  • Strong systems programming (Rust, C++, or similar), with care for code quality and tests

  • Analytical about performance: observe and profile first, form a hypothesis, test it, then change the code and measure again

  • Low ego: ask the naive question, take feedback well, pick up slack outside your job description

  • Enjoy pair programming (we love to pair!) and care about the societal impacts of your work

Preferred Qualifications
  • Experience inside an LLM serving engine and a sense of where its abstractions strain

  • GPU/Accelerator programming

  • OS internals

  • Language modeling with transformers

  • Experience building an allocator, cache, scheduler, or high-bandwidth transport

  • Fluency in Rust

  • Experience making systems reproducible: determinism, replay, property-based tests

The annual compensation range for this role is listed below.

For sales roles, the range provided is the role's On Target Earnings ("OTE&

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Performance Engineer, Inference Engine
Performance Engineer, Inference Engine

Anthropic • San Francisco (CA), New York (NY)

On-site
USD 350,000 - 850,000
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Performance Engineer, Inference Systems San Francisco, CA | New York City, NY | Seattle, WA
Performance Engineer, Inference Systems San Francisco, CA | New York City, NY | Seattle, WA

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Staff+ Software Engineer, Inference Velocity
Staff+ Software Engineer, Inference Velocity

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Staff+ Software Engineer, Inference Velocity
Staff+ Software Engineer, Inference Velocity

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 485,000
Staff+ Software Engineer (Inference Runtime)
Staff+ Software Engineer (Inference Runtime)

Anthropic • San Francisco (CA), Seattle (WA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Staff+ Software Engineer, Inference Runtime
Staff+ Software Engineer, Inference Runtime

Anthropic • New York (NY)

Hybrid
USD 405,000 - 485,000
Generous vacation
Parental leave
Flexible working hours
+1
Performance Engineer, Inference Engine - Flexible Hours
Performance Engineer, Inference Engine - Flexible Hours

Anthropic • San Francisco (CA), New York (NY)

On-site
USD 350,000 - 850,000
Engineering Manager, Accelerator Platform
Engineering Manager, Accelerator Platform

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Staff + Senior Software Engineer, Inference San Francisco, CA | New York City, NY | Seattle, WA
Staff + Senior Software Engineer, Inference San Francisco, CA | New York City, NY | Seattle, WA

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000