ML Systems Engineer: Distributed LLM Training & Inference

Scale AI

Seattle, New York, San Francisco (WA, NY, CA)

On-site

USD 200,800 - 251,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive health coverage
Equity-based compensation
Retirement benefits
Learning and development stipend
Generous PTO
Commuter stipend

Job summary

A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 - $251,000, along with comprehensive benefits.

Qualifications

  • Strong excitement about system optimization.
  • Experience with multi-node LLM training and inference.
  • Strong software engineering skills.

Responsibilities

  • Build, profile and optimize the training and inference framework.
  • Collaborate with ML teams to accelerate research and development.
  • Research and integrate state-of-the-art technologies.

Skills

System optimization
Multi-node LLM training
Large-scale distributed ML systems
CUDA
Pytorch
Transformers
Flash attention
Communication skills

Job description

A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 - $251,000, along with comprehensive benefits.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed ML Systems Engineer for LLM Training
Distributed ML Systems Engineer for LLM Training

Scale AI, Inc. • Seattle (WA)

On-site
USD 189,000 - 237,000
Equity-based compensation
Commuter stipend
Health benefits
ML Systems Engineer — Production-Scale LLM Inference
ML Systems Engineer — Production-Scale LLM Inference

ChipAgents • San Jose (CA)

On-site
USD 150,000 - 350,000
Unlimited PTO
Full benefits (medical, vision, dental, 401k)
Free parking and private gym
Tech Lead for Distributed ML Systems & Training Platform
Tech Lead for Distributed ML Systems & Training Platform

Scale AI • New York (NY), San Francisco (CA)

On-site
USD 275,000 - 350,000
Health, dental, and vision coverage
Retirement benefits
Learning and development stipend
+2
ML Systems Engineer - Scalable Training & Inference
ML Systems Engineer - Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Equity
Benefits
Commuter stipend
ML Infrastructure Engineer for Scalable LLMs
ML Infrastructure Engineer for Scalable LLMs

ServiceNow • Mountain View (CA)

Hybrid
USD 130,000 - 160,000
LLM Inference Architect & Systems Optimizer
LLM Inference Architect & Systems Optimizer

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 230,000
Health insurance
Startup equity
Competitive benefits
Senior ML Systems Engineer: Scalable Training Frameworks
Senior ML Systems Engineer: Scalable Training Frameworks

Cohere • San Francisco (CA)

Hybrid
USD 150,000 - 180,000
ML Engineer — Production-Grade AI & LLM/VLM Systems
ML Engineer — Production-Grade AI & LLM/VLM Systems

Nace.AI • Palo Alto (CA)

On-site
USD 120,000 - 160,000
LLM Training Infra Engineer - Distributed Systems
LLM Training Infra Engineer - Distributed Systems

ByteDance • San Jose (CA)

On-site
USD 244,000 - 450,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+5
Hybrid LLM Engineer: Build & Deploy Scalable AI Models
Hybrid LLM Engineer: Build & Deploy Scalable AI Models

Take2 Consulting, LLC • San Jose (CA)

Hybrid
USD 175,000 - 185,000