Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization across multi‑node GPU clusters.

Join a team focused on scalable AI foundations, monitoring tools, and robust performance improvements for large‑scale runs.

Qualifications

  • 8+ years of distributed systems, ML infrastructure, or high-performance computing experience.
  • Production‑grade Python programming expertise.
  • Low‑level performance mastery: CUDA/cuDNN/Triton and CPU‑GPU interactions.
  • Experience with PyTorch training jobs and multi‑node scaling.
  • Track record tuning hardware‑software interactions for maximum utilization.

Responsibilities

  • Profile bottlenecks across the foundation model training stack and eliminate them.
  • Design, build, and optimize distributed training systems for multi‑node GPU clusters.
  • Implement low‑level code (CUDA/cuDNN/Triton) and integrate into high‑level training frameworks.
  • Optimize workloads for hardware efficiency: CPU/GPU balance, memory, throughput, and networking.
  • Develop monitoring and debugging tools for large‑scale runs to diagnose performance regressions.

Skills

Distributed systems
Python
PyTorch
CUDA
cuDNN
Triton
Kernel optimization
High performance
Data pipelines
GPU utilization

Tools

PyTorch
CUDA
cuDNN
Triton

Job description

Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization across multi‑node GPU clusters.

Join a team focused on scalable AI foundations, monitoring tools, and robust performance improvements for large‑scale runs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer for Distributed GPU Training
Senior ML Infra Engineer for Distributed GPU Training

Genesis Molecular AI • City of Utica (NY)

On-site
USD 150,000 - 190,000
Competitive compensation with salary +
Senior Inference Systems Engineer (GPU/On-Device)
Senior Inference Systems Engineer (GPU/On-Device)

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior ML Systems Engineer – Distributed Training
Senior ML Systems Engineer – Distributed Training

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior HPC Performance Engineer — AI for Science (Equity)
Senior HPC Performance Engineer — AI for Science (Equity)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior ML Systems Engineer – GPU HPC & Distributed Training
Senior ML Systems Engineer – GPU HPC & Distributed Training

Proxima • Boston (MA)

On-site
USD 180,000 - 280,000
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2
Senior Inference Systems Engineer - Low-Latency & On-Device
Senior Inference Systems Engineer - Low-Latency & On-Device

Genesis AI • United States

On-site
USD 180,000 - 240,000
ML Systems Engineer: High-Performance Distributed Training
ML Systems Engineer: High-Performance Distributed Training

Motional • San Francisco (CA)

Hybrid
USD 144,000 - 192,000
Medical
Dental
Vision
+4