Distributed LLM Inference Engineer - Scale HighThroughput AI

Cerebras

Palo Alto (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
Education & Wellbeing Stipend
Paid Parental Leave
Fertility Benefits
Paid Time Off
Commute reimbursement
100% of in-office meals covered

Job summary

Anyscale is seeking a Distributed LLM Inference Engineer in Palo Alto, California. The role focuses on pushing the boundaries of performance for AI inference at large scale, collaborating closely with product teams and open source communities.

The ideal candidate should have experience in running ML inference, familiarity with top deep learning frameworks like PyTorch, and a strong grasp of distributed systems. Attractive benefits and compensation plan included.

Qualifications

  • Familiarity with running ML inference at large scale with high throughput and low latency.
  • Familiarity with deep learning and deep learning frameworks (e.g. PyTorch).
  • Solid understanding of distributed systems and ML inference challenges.

Responsibilities

  • Iterate quickly with product teams to deliver end to end solutions for high scale inference.
  • Integrate Ray Data and LLM engine for optimizations achieving low cost solutions.
  • Contribute to open source software improvements and adoption of techniques.

Skills

Running ML inference at large scale
Deep learning frameworks (e.g., PyTorch)
Distributed systems

Tools

CUDA
vLLM

Job description

Anyscale is seeking a Distributed LLM Inference Engineer in Palo Alto, California. The role focuses on pushing the boundaries of performance for AI inference at large scale, collaborating closely with product teams and open source communities.

The ideal candidate should have experience in running ML inference, familiarity with top deep learning frameworks like PyTorch, and a strong grasp of distributed systems. Attractive benefits and compensation plan included.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed LLM Inference Engineer - Scale AI at Speed
Distributed LLM Inference Engineer - Scale AI at Speed

Anyscale • San Francisco (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans
401k Retirement Plan
+6
Senior AI Inference Engineer — Scale LLMs & PyTorch
Senior AI Inference Engineer — Scale LLMs & PyTorch

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 230,000
Health insurance
Startup equity
Competitive benefits
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Machine Learning Engineer
Machine Learning Engineer

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Staff Software Engineer: LLM Inference & GPU Infra (Equity)
Staff Software Engineer: LLM Inference & GPU Infra (Equity)

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
ML Systems Engineer - Scalable Training & Inference
ML Systems Engineer - Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Equity
Benefits
Commuter stipend
Distributed LLM Inference Engineer
Distributed LLM Inference Engineer

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2