Training / AI Infrastructure

Genesis AI

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization across multi‑node GPU clusters.

Join a team focused on scalable AI foundations, monitoring tools, and robust performance improvements for large‑scale runs.

Qualifications

  • 8+ years of distributed systems, ML infrastructure, or high-performance computing experience.
  • Production‑grade Python programming expertise.
  • Low‑level performance mastery: CUDA/cuDNN/Triton and CPU‑GPU interactions.
  • Experience with PyTorch training jobs and multi‑node scaling.
  • Track record tuning hardware‑software interactions for maximum utilization.

Responsibilities

  • Profile bottlenecks across the foundation model training stack and eliminate them.
  • Design, build, and optimize distributed training systems for multi‑node GPU clusters.
  • Implement low‑level code (CUDA/cuDNN/Triton) and integrate into high‑level training frameworks.
  • Optimize workloads for hardware efficiency: CPU/GPU balance, memory, throughput, and networking.
  • Develop monitoring and debugging tools for large‑scale runs to diagnose performance regressions.

Skills

Distributed systems
Python
PyTorch
CUDA
cuDNN
Triton
Kernel optimization
High performance
Data pipelines
GPU utilization

Tools

PyTorch
CUDA
cuDNN
Triton

Job description

What You’ll Do
  • Drive down wall‑clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack stack, from data pipelines to GPU kernels

  • Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization

  • Implement efficient low‑level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high‑level training frameworks

  • Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networking

  • Develop monitoring and debugging tools for large‑scale runs, enabling rapid diagnosis of performance regressions and failures

What You’ll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years)

  • Production‑grade expertise in Python

  • Low‑level performance mastery: CUDA/cuDNN/Triton, CPU–GPU interactions, data movement, and kernel optimization

  • Scaling at the frontier: experience with PyTorch and training jobs using data, context, pipeline, and model parallelism

  • System‑level mindset with a track record of tuning hardware–software interactions for maximum utilization

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda • California (MO)

On-site
USD 180,000 - 240,000
Senior Principal AI Engineer
Senior Principal AI Engineer

cerence • United States

On-site
USD 180,000 - 250,000
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

TensorScale AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000