ML Systems Engineer: Distributed LLM Training & Inference
Scale AI
Seattle, New York, San Francisco (WA, NY, CA)
On-site
USD 200,800 - 251,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Comprehensive health coverage
Equity-based compensation
Retirement benefits
Learning and development stipend
Generous PTO
Commuter stipend
Job summary
A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 - $251,000, along with comprehensive benefits.
Qualifications
Strong excitement about system optimization.
Experience with multi-node LLM training and inference.
Strong software engineering skills.
Responsibilities
Build, profile and optimize the training and inference framework.
Collaborate with ML teams to accelerate research and development.
Research and integrate state-of-the-art technologies.
Skills
System optimization
Multi-node LLM training
Large-scale distributed ML systems
CUDA
Pytorch
Transformers
Flash attention
Communication skills
Job description
A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 - $251,000, along with comprehensive benefits.