Machine Learning Infrastructure Engineer

David Joseph & Company

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

30 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

David Joseph & Company seeks a talented ML infrastructure engineer to own the distributed training and inference backbone for a large foundation model. You will stand up clusters, build data pipelines for petabyte-scale datasets, and squeeze GPU performance across model scales.

You will work with FSDP/DeepSpeed, NVIDIA GPUs, Linux, Python, C++, Kubernetes/Docker, and major cloud platforms (GCP/AWS/Azure). Relocation supported for San Francisco, on-site five days weekly.

Qualifications

  • 2–10 years building ML infra for foundation models.
  • Experience with large-scale distributed training from scratch.
  • Background in science-focused or physical-AI domains is a plus.

Responsibilities

  • Design, deploy, and maintain large distributed ML training and inference clusters.
  • Build end-to-end pipelines to manage petabyte-scale datasets and training.
  • Research training approaches and parallelization techniques.
  • Profile and optimize low-level GPU operations for performance.
  • Track research developments and apply new ideas.

Skills

Distributed training
Large-scale infra
Linux
Python
C++
Kubernetes
Docker
Cloud platforms

Tools

FSDP
DeepSpeed
CUDA
NVIDIA GPUs

Job description

  • Full-time Compensation: $200K–$400K + competitive early-stage equity

San Francisco, CA

  • On-site (5 days/week)
  • Full-time Compensation: $200K–$400K + competitive early-stage equity
About The Company

Our client is a Series A AI research lab building large-scale foundation models for scientific and physical-AI domains. Backed by top-tier investors, they are pursuing a deliberately non-consensus technical thesis and are among the best-funded teams in their space. The founding team comes from self-driving, robotics, and scientific research, and they are scaling their research and engineering org significantly this year.

Founded 2024

  • Small, fast-growing team
  • Industry: AI / foundation models / physical AI
The Role

You would own the distributed training and inference backbone for a foundation model trained from scratch — standing up clusters, building data and training pipelines at petabyte scale, and squeezing performance out of GPUs at a low level across model scales.

What You'll Be Doing
  • Design, deploy, and maintain large distributed ML training and inference clusters
  • Build efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and training across the full ML lifecycle
  • Research and test training approaches, including parallelization techniques and numerical-precision trade-offs across model scales
  • Profile and debug low-level GPU operations to optimize performance
  • Track new research and bring fresh ideas into the work

Tech stack: Distributed training frameworks (FSDP, DeepSpeed), NVIDIA GPUs, Linux, Python, C++, Kubernetes/Docker, and a major cloud platform (GCP, AWS, or Azure).

Requirements
  • 2–10 years building large-scale ML infrastructure for core foundation models
  • Hands-on experience building infrastructure for foundation models trained from scratch, rather than fine-tuning existing models
  • A background at a science-focused or physical-AI company (for example self-driving, robotics, or biology)
  • Deep, demonstrable expertise optimizing large-scale training and inference workloads
  • Working proficiency with distributed training frameworks such as FSDP or DeepSpeed
  • A clear pattern of intentional, mission-driven career decisions
  • Able to work on-site 5 days/week in San Francisco (relocation supported)
Nice to Haves
  • Generalist experience spanning the full ML lifecycle
  • Low-level GPU performance optimization and debugging (CUDA, JAX)
Why Join
  • Take a bet on a distinctive, non-consensus approach to building intelligence
  • Join early, with real ownership of the training and inference backbone
  • Work in a domain with fast, objective ground-truth feedback and data at a scale beyond typical LLM training
  • Well-funded and building a strong, senior research and engineering team
Details
  • Location: San Francisco, CA
  • Work policy: In-person 5 days/week (relocation supported)
  • Compensation: $200K–$400K + competitive early-stage equity
  • Visa sponsorship: Open to supporting work authorization for the right candidate
  • Employment type: Full-time
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Causal Labs — Machine Learning Infrastructure Engineer
Causal Labs — Machine Learning Infrastructure Engineer

davidjoseph-co • San Francisco (CA)

On-site
USD 200,000 - 400,000
Relocation support
Visa sponsorship
On-site in SF
+1
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Recruiting From Scratch • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Competitive equity
Cutting-edge diffusion models
Direct collaboration with researchers
Research Engineer Infrastructure Training Systems
Research Engineer Infrastructure Training Systems

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health benefits
Unlimited PTO
Paid parental leave
+1
SF ML Infrastructure Architect – Foundation Models
SF ML Infrastructure Architect – Foundation Models

davidjoseph-co • San Francisco (CA)

On-site
USD 200,000 - 400,000
Relocation support
Visa sponsorship
On-site in SF
+1
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Founding Engineer, SWE
Founding Engineer, SWE

David Joseph & Company • San Francisco (CA)

On-site
USD 150,000 - 200,000
AI Training Infrastructure Engineer
AI Training Infrastructure Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3