SF ML Infrastructure Architect – Foundation Models

davidjoseph-co

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation support
Visa sponsorship
On-site in SF
Equity compensation

Job summary

Unknown Company in San Francisco, CA seeks a lead ML infrastructure engineer to own the distributed training and inference backbone for a foundation model trained from scratch. You will stand up clusters, build data pipelines at petabyte scale, and optimize GPU performance across model scales.

You will work with FSDP/DeepSpeed, NVIDIA GPUs, Linux, Python and C++, in a distributed cloud environment across GCP/AWS/Azure, with relocation supported.

Qualifications

  • 2–10 years building large-scale ML infrastructure for core foundation models.
  • Hands-on experience building infrastructure for foundation models trained from scratch.
  • Background at science-focused or physical-AI company (e.g., robotics, biology).
  • Deep, demonstrable expertise optimizing large-scale training and inference workloads.
  • Experience with distributed training frameworks such as FSDP or DeepSpeed.
  • Able to relocate and work on-site in San Francisco 5 days/week.

Responsibilities

  • Design, deploy, and maintain large distributed ML training and inference clusters.
  • Build scalable end-to-end pipelines for petabyte-scale datasets and training.
  • Research training approaches and parallelization techniques.
  • Profile and debug low-level GPU operations to optimize performance.
  • Stay updated with new research and bring innovative ideas.

Skills

Distributed training
GPU optimization
Python
C++
Cloud platforms

Tools

FSDP
DeepSpeed
NVIDIA GPUs
Linux
Kubernetes
Docker
GCP
AWS
Azure

Job description

Unknown Company in San Francisco, CA seeks a lead ML infrastructure engineer to own the distributed training and inference backbone for a foundation model trained from scratch. You will stand up clusters, build data pipelines at petabyte scale, and optimize GPU performance across model scales.

You will work with FSDP/DeepSpeed, NVIDIA GPUs, Linux, Python and C++, in a distributed cloud environment across GCP/AWS/Azure, with relocation supported.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer for Petabyte-Scale Training (SF On-Site)
ML Infra Engineer for Petabyte-Scale Training (SF On-Site)

David Joseph & Company • San Francisco (CA)

On-site
USD 200,000 - 400,000
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

David Joseph & Company • San Francisco (CA)

On-site
USD 200,000 - 400,000
Causal Labs — Machine Learning Infrastructure Engineer
Causal Labs — Machine Learning Infrastructure Engineer

davidjoseph-co • San Francisco (CA)

On-site
USD 200,000 - 400,000
Relocation support
Visa sponsorship
On-site in SF
+1
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Staff Foundation Model Inference Engineer
Staff Foundation Model Inference Engineer

United States Digital Space LLC • San Francisco (CA)

On-site
USD 190,000 - 265,000
Senior ML Engineer: Foundation Models & Scalable Systems
Senior ML Engineer: Foundation Models & Scalable Systems

AI Breaking Wire • Mountain View (CA)

On-site
USD 200,000 - 350,000
Stock grants
Health/dental/vision benefits
Paid time off & parental leave
+1
Senior Foundation Model Systems Engineer
Senior Foundation Model Systems Engineer

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Foundation Models AI/ML Engineer – Systems & GPU Optimized
Foundation Models AI/ML Engineer – Systems & GPU Optimized

Amazon Science • Santa Clara (CA)

On-site
USD 172,000 - 222,000
Senior ML Infra Architect: Scalable Data Pipelines
Senior ML Infra Architect: Scalable Data Pipelines

Adobe Inc. • San Jose (CA)

On-site
USD 268,000 - 388,000
Senior ML Infra Engineer - Large-Scale Training & Pipelines
Senior ML Infra Engineer - Large-Scale Training & Pipelines

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000