Distributed LLM Inference Engineer - Scale AI at Speed

Anyscale

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock Options
Healthcare plans
401k Retirement Plan
Education & Wellbeing Stipend
Paid Parental Leave
Fertility Benefits
Paid Time Off
Commute reimbursement
100% of in-office meals covered

Job summary

Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open-source technologies and contributing to community projects. Candidates should have a solid understanding of distributed systems and familiarity with deep learning frameworks, ideally with experience in PyTorch and Ray. Anyscale offers competitive compensation and extensive benefits, including healthcare coverage and stock options.

Qualifications

  • Familiarity with running ML inference at large scale with high throughput and low latency.
  • Familiarity with deep learning and deep learning frameworks (e.g. PyTorch).
  • Solid understanding of distributed systems and ML inference challenges.

Responsibilities

  • Iterate with product teams to ship end-to-end solutions for inference.
  • Work across the stack integrating Ray Data and LLM engine for optimizations.
  • Contribute improvements to open source software like vLLM.

Skills

ML inference at large scale
Deep learning frameworks
Distributed systems knowledge

Tools

PyTorch
Ray
CUDA

Job description

Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open-source technologies and contributing to community projects. Candidates should have a solid understanding of distributed systems and familiarity with deep learning frameworks, ideally with experience in PyTorch and Ray. Anyscale offers competitive compensation and extensive benefits, including healthcare coverage and stock options.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed LLM Inference Engineer - Scale HighThroughput AI
Distributed LLM Inference Engineer - Scale HighThroughput AI

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6
Distributed LLM Inference Engineer
Distributed LLM Inference Engineer

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6
Distributed LLM Inference Engineer
Distributed LLM Inference Engineer

Anyscale • San Francisco (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans
401k Retirement Plan
+6
Senior AI Inference Engineer — Scale LLMs & PyTorch
Senior AI Inference Engineer — Scale LLMs & PyTorch

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 230,000
Health insurance
Startup equity
Competitive benefits
Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
ML Systems Engineer: Distributed LLM Training & Inference
ML Systems Engineer: Distributed LLM Training & Inference

Scale AI • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 200,000 - 251,000
Comprehensive health coverage
Equity-based compensation
Retirement benefits
+3
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
Backend Engineer – AI Infra for Scalable LLMs (SF)
Backend Engineer – AI Infra for Scalable LLMs (SF)

Anara • San Francisco (CA)

On-site
USD 150,000 - 200,000
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Lead Engineer, Inference Platform & Scale
Lead Engineer, Inference Platform & Scale

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Job stability with startup vitality
Open-source AI research
Simple, non-corporate work culture