Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization across multi‑node GPU clusters.

Join a team focused on scalable AI foundations, monitoring tools, and robust performance improvements for large‑scale runs.

Qualifications

  • 8+ years of distributed systems, ML infrastructure, or high-performance computing experience.
  • Production‑grade Python programming expertise.
  • Low‑level performance mastery: CUDA/cuDNN/Triton and CPU‑GPU interactions.
  • Experience with PyTorch training jobs and multi‑node scaling.
  • Track record tuning hardware‑software interactions for maximum utilization.

Responsibilities

  • Profile bottlenecks across the foundation model training stack and eliminate them.
  • Design, build, and optimize distributed training systems for multi‑node GPU clusters.
  • Implement low‑level code (CUDA/cuDNN/Triton) and integrate into high‑level training frameworks.
  • Optimize workloads for hardware efficiency: CPU/GPU balance, memory, throughput, and networking.
  • Develop monitoring and debugging tools for large‑scale runs to diagnose performance regressions.

Skills

Distributed systems
Python
PyTorch
CUDA
cuDNN
Triton
Kernel optimization
High performance
Data pipelines
GPU utilization

Tools

PyTorch
CUDA
cuDNN
Triton

Job description

Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization across multi‑node GPU clusters.

Join a team focused on scalable AI foundations, monitoring tools, and robust performance improvements for large‑scale runs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer for Distributed GPU Training
Senior ML Infra Engineer for Distributed GPU Training

Genesis Molecular AI • City of Utica (NY)

On-site
USD 150,000 - 190,000
Competitive compensation with salary +
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
GPU Systems Engineer — Distributed Training & Inference
GPU Systems Engineer — Distributed Training & Inference

TensorScale AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Training Systems Architect (Distributed)
AI Training Systems Architect (Distributed)

Unconventional AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Health benefits
401k matching
Unlimited PTO
+1
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior AI Distributed Systems Engineer
Senior AI Distributed Systems Engineer

United States Digital Space LLC • United States

Remote
USD 180,000 - 240,000
Lead ML Systems Engineer — Distributed GPU Training & Infra
Lead ML Systems Engineer — Distributed GPU Training & Infra

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Senior AI Systems Architect: Scalable Distributed Training
Senior AI Systems Architect: Scalable Distributed Training

Cerence AI • United States

On-site
USD 180,000 - 240,000
Distributed AI Training Architect
Distributed AI Training Architect

cerence • United States

On-site
USD 180,000 - 250,000
Distributed ML Research Scientist — Frontier-Scale Training
Distributed ML Research Scientist — Frontier-Scale Training

Ifm Us • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Health benefits
401K Plan
Paid time off
+2