ML Performance Engineer: Distributed Training & Inference

Veeda AI

Seattle (WA)

On-site

USD 180,000 - 260,000

Full time

19 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Veeda AI is seeking a highly skilled engineer to advance distributed training and high-performance inference for multimodal world models. You will own multi-node throughput, profile GPU utilization, and optimize kernels and CUDA/Triton code across NVLink fabrics.

Candidates should have deep PyTorch expertise, Python and C++/CUDA fluency, and experience with frameworks like FSDP2, Megatron-Core, DeepSpeed. You will work in a fast-moving research environment to push physical AI capabilities

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or related field.
  • Experience with PyTorch and a large-scale parallelism stack (FSDP2, Megatron-Core, TorchTitan, or DeepSpeed).
  • Fluency in Python and C++/CUDA with the ability to predict kernel stalls before profiling.
  • Experience profiling live training/inference with Nsight Systems or PyTorch profiler.
  • Expertise in low-precision numerics, kernel authoring, or fault diagnosis, with credibility in others.
  • Experience optimizing training and inference on very large GPU clusters, including topology-aware placement and scaling efficiency.

Responsibilities

  • Own step time and model FLOPs utilization for multi-node video world model training, selecting parallelism strategies.
  • Profile end-to-end latency, GPU utilization, memory pressure, and kernel efficiency; improve throughput with advanced techniques.
  • Address precision and numerical stability challenges across FP16/FP8 and related formats in video tokenizer and VAE layers.
  • Write and tune CUDA and Triton kernels to optimize attention and kernel performance beyond standard PyTorch implementations.
  • Tune NCCL collectives and overlap across NVLink domains to maximize throughput.
  • Develop fault diagnostics and asynchronous checkpointing to minimize interruption costs.

Skills

Python
C++
CUDA
PyTorch
Distributed training
Performance profiling

Education

Bachelor's degree in Computer Science or related field

Tools

PyTorch FSDP2
Megatron-Core
TorchTitan
DeepSpeed
Nsight Systems
NCCL

Job description

Veeda AI is seeking a highly skilled engineer to advance distributed training and high-performance inference for multimodal world models. You will own multi-node throughput, profile GPU utilization, and optimize kernels and CUDA/Triton code across NVLink fabrics.

Candidates should have deep PyTorch expertise, Python and C++/CUDA fluency, and experience with frameworks like FSDP2, Megatron-Core, DeepSpeed. You will work in a fast-moving research environment to push physical AI capabilities

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Performance Engineer – Multi-Node Training & Kernels
ML Performance Engineer – Multi-Node Training & Kernels

Veeda AI • California

On-site
USD 180,000 - 240,000
ML Systems Engineer — Performance & Scale
ML Systems Engineer — Performance & Scale

Veeda Innovation • Northern (KY)

Hybrid
USD 150,000 - 230,000
Staff ML Engineer — Distributed AI Systems & Tools
Staff ML Engineer — Distributed AI Systems & Tools

Veeda AI • California

On-site
USD 120,000 - 190,000
AI Infra Engineer: GPU Clusters & HPC
AI Infra Engineer: GPU Clusters & HPC

Veeda AI • California

On-site
USD 150,000 - 210,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda Innovation • Northern (KY)

On-site
USD 150,000 - 230,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
ML Data Engineer, Multimodal AI & Data Pipelines
ML Data Engineer, Multimodal AI & Data Pipelines

Veeda AI • California

On-site
USD 140,000 - 210,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Senior DL Engineer – Inference & Model Optimization (Equity)
Senior DL Engineer – Inference & Model Optimization (Equity)

NVIDIA • United States

Remote
USD 184,000 - 356,000
Equity
Benefits
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000