Distributed LLM Inference Engineer - Scale AI at Speed
Anyscale
San Francisco (CA)
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Stock Options
Healthcare plans
401k Retirement Plan
Education & Wellbeing Stipend
Paid Parental Leave
Fertility Benefits
Paid Time Off
Commute reimbursement
100% of in-office meals covered
Job summary
Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open-source technologies and contributing to community projects. Candidates should have a solid understanding of distributed systems and familiarity with deep learning frameworks, ideally with experience in PyTorch and Ray. Anyscale offers competitive compensation and extensive benefits, including healthcare coverage and stock options.
Qualifications
Familiarity with running ML inference at large scale with high throughput and low latency.
Familiarity with deep learning and deep learning frameworks (e.g. PyTorch).
Solid understanding of distributed systems and ML inference challenges.
Responsibilities
Iterate with product teams to ship end-to-end solutions for inference.
Work across the stack integrating Ray Data and LLM engine for optimizations.
Contribute improvements to open source software like vLLM.
Skills
ML inference at large scale
Deep learning frameworks
Distributed systems knowledge
Tools
PyTorch
Ray
CUDA
Job description
Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open-source technologies and contributing to community projects. Candidates should have a solid understanding of distributed systems and familiarity with deep learning frameworks, ideally with experience in PyTorch and Ray. Anyscale offers competitive compensation and extensive benefits, including healthcare coverage and stock options.