LLM Inference Engineer — Distributed Systems

OpenTalent

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco to push the performance boundaries of large-scale inference. You will collaborate with product teams to deliver end-to-end batch and online inference solutions, leveraging Ray Data and LLM engines.

The role requires familiarity with deep learning frameworks (PyTorch), distributed systems, and GPUs/CUDA, with opportunities to contribute to open source projects like vLLM and TensorRT-LLM.

Qualifications

  • Familiarity with running ML inference at large scale with high throughput and low latency.
  • Familiarity with deep learning and deep learning frameworks (e.g. PyTorch).
  • Solid understanding of distributed systems, ML inference challenges.

Responsibilities

  • Iterate quickly with product teams to ship end-to-end batch and online inference solutions at high scale.
  • Work across the stack integrating Ray Data and LLM engine providing optimizations for low-cost large-scale ML inference.
  • Integrate with open source software like vLLM and contribute improvements to open source.
  • Follow state-of-the-art in open source and research, implementing best practices.

Skills

ML inference at scale
Deep learning
PyTorch familiarity
Distributed systems

Tools

Ray
vLLM
TensorRT-LLM
CUDA

Job description

Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco to push the performance boundaries of large-scale inference. You will collaborate with product teams to deliver end-to-end batch and online inference solutions, leveraging Ray Data and LLM engines.

The role requires familiarity with deep learning frameworks (PyTorch), distributed systems, and GPUs/CUDA, with opportunities to contribute to open source projects like vLLM and TensorRT-LLM.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed LLM Inference Engineer - Scale HighThroughput AI
Distributed LLM Inference Engineer - Scale HighThroughput AI

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6
Distributed LLM Inference Engineer
Distributed LLM Inference Engineer

OpenTalent • San Francisco (CA)

On-site
USD 180,000 - 260,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Distributed LLM Inference Engineer - Scale & Resilience
Distributed LLM Inference Engineer - Scale & Resilience

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Equity
Health coverage
Flexible PTO
+4
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Distributed LLM Inference Engineer
Distributed LLM Inference Engineer

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6
Staff Software Engineer: LLM Inference & GPU Infra (Equity)
Staff Software Engineer: LLM Inference & GPU Infra (Equity)

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
Member of the Technical Staff- LLMs
Member of the Technical Staff- LLMs

Amadeus Search • San Francisco (CA)

Hybrid
USD 170,000 - 220,000
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4