ML Systems Engineer — Performance & Scale

Veeda Innovation

Northern (KY)

Hybrid

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veeda AI is seeking a Member of Technical Staff - ML Performance to optimize distributed training and inference for large-scale video models. You will drive throughput, profile bottlenecks, and implement kernel-level enhancements across PyTorch, CUDA, and specialized parallelism stacks.

The role requires deep experience with FSDP2, Megatron-Core, TorchTitan, or DeepSpeed, plus strong Python/C++ skills. Location in the US is expected, with a focus on scalable acceleration across GPU clusters.

Qualifications

  • Experience with PyTorch and large-scale multi-node parallelism stacks on real workloads.
  • Fluency in Python and C++/CUDA, with kernel performance intuition.
  • Proven ability to profile training/inference and translate traces into gains.

Responsibilities

  • Own training throughput and mix parallelism for multi-node video world model training.
  • Profile latency, GPU utilization, memory pressure, and improve throughput with optimizations.
  • Tune kernels and integrations to reduce step time and enable efficient attention.
  • Optimize communication and overlap across NVLink and NCCL for scalability.
  • Develop fault diagnostics and asynchronous checkpointing to reduce downtime.

Skills

PyTorch
Python
C++/CUDA
FSDP2
Megatron-Core
DeepSpeed
TorchTitan
Profiling

Education

Bachelor's degree in Computer Science or related field

Tools

Nsight Systems
PyTorch Profiler
CUDA
NCCL
torch.compile
Triton
CuTe-DSL

Job description

Veeda AI is seeking a Member of Technical Staff - ML Performance to optimize distributed training and inference for large-scale video models. You will drive throughput, profile bottlenecks, and implement kernel-level enhancements across PyTorch, CUDA, and specialized parallelism stacks.

The role requires deep experience with FSDP2, Megatron-Core, TorchTitan, or DeepSpeed, plus strong Python/C++ skills. Location in the US is expected, with a focus on scalable acceleration across GPU clusters.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Performance Engineer — Distributed Training & Kernels
ML Performance Engineer — Distributed Training & Kernels

Veeda • California (MO)

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda • California (MO)

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda Innovation • Northern (KY)

Hybrid
USD 150,000 - 230,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
AI Infrastructure Architect for HPC & GPU Clusters
AI Infrastructure Architect for HPC & GPU Clusters

Veeda • California (MO)

On-site
USD 150,000 - 200,000
Staff Data Engineer, Multimodal AI Pipelines
Staff Data Engineer, Multimodal AI Pipelines

Veeda • California (MO)

On-site
USD 130,000 - 180,000
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Technical Staff Engineer - World Models & Multimodal AI
Technical Staff Engineer - World Models & Multimodal AI

Veeda • California (MO)

On-site
USD 130,000 - 200,000
Senior ML Performance Engineer — Ultra-Fast Inference + Equity
Senior ML Performance Engineer — Ultra-Fast Inference + Equity

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
GPU ML Infra Intern: Speed Up Training & Profiling
GPU ML Infra Intern: Speed Up Training & Profiling

Plus 2 • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Free lunch
Snacks & drinks
401(k) plan