ML Performance Engineer – Multi-Node Training & Kernels

Veeda AI

California

On-site

USD 180,000 - 240,000

Full time

33 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Veeda AI in California is building the next generation of multimodal foundation world models for Physical AI. You will design and optimize distributed training and inference pipelines on large GPU clusters, focusing on low-latency, high-throughput workloads.

The role requires deep expertise in PyTorch, large-scale parallelism, and CUDA kernel work, with a track record of improving hardware utilization and stability at scale.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or related field.
  • Deep hands-on experience with PyTorch and large-scale parallelism stacks on real multi-node jobs.
  • Fluency in Python and C++/CUDA with kernel profiling intuition.
  • Experience profiling live training/inference with Nsight Systems or PyTorch profiler and translating traces into improvements.
  • Expertise in low-precision numerics, kernel authoring, or large-run fault diagnosis.

Responsibilities

  • Distributed Training Throughput: own step time and model FLOPs for multi-node video world model training.
  • Inference Throughput Optimization: profile latency, GPU utilization, memory pressure, and kernel efficiency; improve throughput via profiling and CUDA tools.
  • Precision & Numerical Stability: work with BF16, FP8, and related recipes to ensure stability in video tokenizer and VAE layers.
  • Kernels & Compilation: write/tune CUDA and Triton kernels; drive FlashAttention-4 and related integration to reduce step time.
  • Communication & Overlap: optimize NCCL collectives and compute/communication overlap across NVLink.
  • Fault Diagnostics & Recovery: build detection for silent data corruption and asynchronous checkpointing.

Skills

PyTorch
Python
C++/CUDA
FSDP2
Megatron-Core
DeepSpeed
NCCL
Nsight Systems
Profiling

Education

Bachelor's degree in CS/CE

Tools

NCCL
Nsight Systems
CUDA
PyTorch Profiler

Job description

Veeda AI in California is building the next generation of multimodal foundation world models for Physical AI. You will design and optimize distributed training and inference pipelines on large GPU clusters, focusing on low-latency, high-throughput workloads.

The role requires deep expertise in PyTorch, large-scale parallelism, and CUDA kernel work, with a track record of improving hardware utilization and stability at scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Performance Engineer: Distributed Training & Inference
ML Performance Engineer: Distributed Training & Inference

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 260,000
ML Systems Engineer — Performance & Scale
ML Systems Engineer — Performance & Scale

Veeda Innovation • Northern (KY)

Hybrid
USD 150,000 - 230,000
AI Infra Engineer: GPU Clusters & HPC
AI Infra Engineer: GPU Clusters & HPC

Veeda AI • California

On-site
USD 150,000 - 210,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda Innovation • Northern (KY)

On-site
USD 150,000 - 230,000
Staff ML Engineer — Distributed AI Systems & Tools
Staff ML Engineer — Distributed AI Systems & Tools

Veeda AI • California

On-site
USD 120,000 - 190,000
ML Data Engineer, Multimodal AI & Data Pipelines
ML Data Engineer, Multimodal AI & Data Pipelines

Veeda AI • California

On-site
USD 140,000 - 210,000
Performance Engineer - ML Training & CUDA Kernels
Performance Engineer - ML Training & CUDA Kernels

Cohere • New York (NY)

Hybrid
USD 150,000 - 190,000
Lunch stipend
Health benefits
RRSP matching
+5
Distributed Training Engineer for Multimodal Models
Distributed Training Engineer for Multimodal Models

Luma • Redwood City (CA)

On-site
USD 210,000 - 260,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Senior Distributed AI Training Architect
Senior Distributed AI Training Architect

Luma AI • United States

Remote
USD 180,000 - 280,000