Member of Technical Staff - ML Performance

Veeda AI

Seattle (WA)

On-site

USD 180,000 - 260,000

Full time

16 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Veeda AI is seeking a highly skilled engineer to advance distributed training and high-performance inference for multimodal world models. You will own multi-node throughput, profile GPU utilization, and optimize kernels and CUDA/Triton code across NVLink fabrics.

Candidates should have deep PyTorch expertise, Python and C++/CUDA fluency, and experience with frameworks like FSDP2, Megatron-Core, DeepSpeed. You will work in a fast-moving research environment to push physical AI capabilities

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or related field.
  • Experience with PyTorch and a large-scale parallelism stack (FSDP2, Megatron-Core, TorchTitan, or DeepSpeed).
  • Fluency in Python and C++/CUDA with the ability to predict kernel stalls before profiling.
  • Experience profiling live training/inference with Nsight Systems or PyTorch profiler.
  • Expertise in low-precision numerics, kernel authoring, or fault diagnosis, with credibility in others.
  • Experience optimizing training and inference on very large GPU clusters, including topology-aware placement and scaling efficiency.

Responsibilities

  • Own step time and model FLOPs utilization for multi-node video world model training, selecting parallelism strategies.
  • Profile end-to-end latency, GPU utilization, memory pressure, and kernel efficiency; improve throughput with advanced techniques.
  • Address precision and numerical stability challenges across FP16/FP8 and related formats in video tokenizer and VAE layers.
  • Write and tune CUDA and Triton kernels to optimize attention and kernel performance beyond standard PyTorch implementations.
  • Tune NCCL collectives and overlap across NVLink domains to maximize throughput.
  • Develop fault diagnostics and asynchronous checkpointing to minimize interruption costs.

Skills

Python
C++
CUDA
PyTorch
Distributed training
Performance profiling

Education

Bachelor's degree in Computer Science or related field

Tools

PyTorch FSDP2
Megatron-Core
TorchTitan
DeepSpeed
Nsight Systems
NCCL

Job description

About Us

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.


Responsibilities


  • Distributed Training Throughput: Own step time and model FLOPs utilization for multi-node video world model training, choosing the tensor, context, and expert parallelism mix in PyTorch FSDP2 and Megatron-Core rather than inheriting a default.

  • Inference Throughput Optimization: Profile end-to-end latency, GPU utilization, memory pressure, and kernel efficiency with PyTorch Profiler, Nsight Systems, Nsight Compute, and torch.utils.benchmark, then improve throughput through torch.compile/Inductor, CUDA Graphs, mixed and low precision, quantization, operator fusion, and multi-GPU serving.

  • Precision & Numerical Stability: Take BF16, FP8, and NVFP4 recipes from running to converging on Blackwell, chasing scaling-factor and accumulation bugs into the video tokenizer and VAE layers where the activation outliers actually live.

  • Kernels & Compilation: Write and tune the CUDA and Triton kernels PyTorch does not give us, driving FlashAttention-4, FlexAttention, and torch.compile integration so quadratic attention over long video sequences stops setting step time.

  • Communication & Overlap: Tune NCCL collectives and compute/communication overlap across NVLink domains and the fabric, using the NCCL flight recorder to turn a watchdog timeout into a named rank and collective, not a restart.

  • Fault Diagnostics & Recovery: Build the detection layer for silent data corruption (SDC), stuck CUDA kernels, and "card-freeze" hangs, plus asynchronous and tiered checkpointing that makes an interruption cost minutes rather than a day.


Requirements


  • Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.

  • Deep hands-on experience with PyTorch and at least one large-scale parallelism stack (FSDP2, Megatron-Core, TorchTitan, or DeepSpeed) on real multi-node jobs, not single-node approximations.

  • Fluency in Python and C++/CUDA with the ability to predict where a kernel will stall from its memory access pattern before profiling.

  • Experience profiling live training and inference runs with Nsight Systems or the PyTorch profiler and translating traces into quantifiable step-time or MFU improvements.

  • Expertise in at least one of low-precision numerics, kernel authoring, or large-run fault diagnosis, and credibility in the others.


Nice to Have


  • Experience optimizing training or inference workloads across very large GPU clusters, including topology-aware placement, scaling efficiency, performance isolation, and diagnosing failures that emerge only at fleet scale.

  • Experience writing Triton, CUTLASS, or CuTe-DSL kernels, or contributing to open-source kernel libraries.

  • Experience implementing context or sequence parallelism for long-horizon video or high-token-count models.

  • Experience running or porting large training workloads on AMD GPUs (ROCm) or Google TPUs (JAX/XLA).

  • Experience building fault-tolerant training with elastic world size, dynamic node re-queueing, or asynchronous distributed checkpointing.

  • Experience optimizing generative inference for interactive rollouts, including few-step samplers, distillation, and KV or latent caching.

  • Publications or presentations on machine learning systems, compilers, or high-performance kernels.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda AI • California

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda Innovation • Northern (KY)

On-site
USD 150,000 - 230,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Operations
Member of Technical Staff - ML Operations

Veeda AI • California

On-site
USD 120,000 - 190,000
Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

SF Tensor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Simulation
Member of Technical Staff - Simulation

Veeda AI • California

On-site
USD 120,000 - 180,000
Member of the Technical Staff - Systems ML Engineer
Member of the Technical Staff - Systems ML Engineer

Breakout Ventures • Cambridge (MA)

On-site
USD 180,000 - 270,000
Equity
Lunch subsidy
Health insurance
+1