Member of Technical Staff - ML Infra

Kindredventures

San Francisco (CA)

On-site

USD 160,000 - 220,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Kindredventures is seeking an experienced ML infra/engineering professional to design, deploy, and maintain large distributed ML training and inference clusters in a production environment.

You will build scalable pipelines for petabyte-scale data and model training, explore parallelization and numeric precision trade-offs across varying model scales, and optimize GPU performance through profiling and low-level debugging.

Qualifications

  • Experience with large distributed ML training and inference systems.
  • Experience with cloud ML services on major clouds (GCP/AWS/Azure).
  • Experience with containers and orchestration (Kubernetes/Docker).
  • Experience with distributed task management and scalable model serving/deployment.
  • Knowledge of monitoring, logging, observability, and version control for ML systems.

Responsibilities

  • Design, deploy, and maintain large distributed ML training and inference clusters
  • Develop scalable end-to-end pipelines to manage petabyte-scale datasets and model training throughout the ML lifecycle
  • Research and test training approaches including parallelization and precision trade-offs across model scales
  • Analyze, profile and debug low-level GPU operations to optimize performance
  • Stay up-to-date on research to bring new ideas to work

Tools

Distributed training frameworks
Cloud platforms
Kubernetes
Docker
Distributed task management
Observability & version control

Job description

Responsibilities
  • Design, deploy, and maintain large distributed ML training and inference clusters
  • Develop efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and model training throughout the entire ML lifecycle
  • Research and test various training approaches including parallelization techniques and numerical precision trade-offs across different model scales
  • Analyze, profile and debug low-level GPU operations to optimize performance
  • Stay up-to-date on research to bring new ideas to work
What we’re looking for

We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.

  • Strong grasp of state-of-the‑art techniques for optimizing training and inference workloads
  • Demonstrated proficiency with distributed training frameworks (e.g. FSDP, DeepSpeed) to train large foundation models
  • Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings
  • Familiarity with containerization and orchestration frameworks (e.g., Kubernetes, Docker)
  • Background working on distributed task management systems and scalable model serving & deployment architectures
  • Understanding of monitoring, logging, observability, and version control best practices for ML systems

You don’t have to meet every single requirement above.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer - Large-Scale Training & Pipelines
Senior ML Infra Engineer - Large-Scale Training & Pipelines

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Member of Technical Staff - ML Research
Member of Technical Staff - ML Research

Kindredventures • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff
Member of Technical Staff

Fireworks AI • United States

On-site
USD 140,000 - 220,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000