ML Performance Engineer - Distributed Training & Inference

Veeda AI

Zürich

On-site

CHF 120,000 - 160,000

Full time

26 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Veeda AI in Zürich, Switzerland is building the next generation of multimodal world models for Physical AI. The role focuses on optimizing multi-node video world model training and inference, managing tensor, context, and expert parallelism across PyTorch FSDP2 and Megatron-Core, not default options.

You will profile end-to-end latency, GPU utilization, memory pressure, and kernel efficiency using PyTorch Profiler, Nsight Systems, and torch.utils.benchmark, then improve throughput via

Qualifications

  • Hands-on experience with PyTorch and large-scale parallelism stacks on multi-node clusters.

Responsibilities

  • Own step time and model FLOPs utilization for multi-node video world model training.
  • Profile end-to-end latency, GPU utilization, memory pressure, and kernel efficiency for inference throughput.
  • Tuning CUDA and Triton kernels to improve throughput and scaling across GPUs.
  • Develop and integrate fault-tolerant and asynchronous checkpointing.
  • Tune NCCL collectives and overlap across NVLink domains.
  • Diagnose silent data corruption and stuck kernels, enabling fast recovery.

Skills

Python
C++/CUDA
PyTorch
FSDP2
Megatron-Core
TorchTitan
DeepSpeed
Nsight Systems
torch.compile
CUDA Graphs

Education

Bachelor's degree in Computer Science/Engineering or related field

Tools

Nsight Systems
PyTorch Profiler
Triton
CuTe-DSL
Torch.compile
CUDA

Job description

Veeda AI in Zürich, Switzerland is building the next generation of multimodal world models for Physical AI. The role focuses on optimizing multi-node video world model training and inference, managing tensor, context, and expert parallelism across PyTorch FSDP2 and Megatron-Core, not default options.

You will profile end-to-end latency, GPU utilization, memory pressure, and kernel efficiency using PyTorch Profiler, Nsight Systems, and torch.utils.benchmark, then improve throughput via

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Scientist, World Models & Multimodal AI
Staff Scientist, World Models & Multimodal AI

Veeda AI • Zürich

On-site
CHF 120,000 - 180,000
Staff ML Engineer: Systems & Experimentation
Staff ML Engineer: Systems & Experimentation

Veeda AI • Zürich

On-site
CHF 140,000 - 190,000
Staff AI Simulation Engineer
Staff AI Simulation Engineer

Veeda AI • Zürich

On-site
CHF 120,000 - 180,000
Member of Technical Staff - World Models
Member of Technical Staff - World Models

Veeda AI • Zürich

On-site
CHF 120,000 - 180,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda AI • Zürich

On-site
CHF 120,000 - 160,000
Member of Technical Staff - ML Data
Member of Technical Staff - ML Data

Veeda AI • Zürich

On-site
CHF 120,000 - 180,000
Senior ML Engineer, Inference & Optimization
Senior ML Engineer, Inference & Optimization

Nebius Group • Zürich

On-site
CHF 150,000 - 220,000
Competitive compensation
Career growth and learning机会
Flexibility and ownership
+3
AI Infra Engineer — GPU Clusters & HPC
AI Infra Engineer — GPU Clusters & HPC

Veeda AI • Zürich

On-site
CHF 120,000 - 180,000
AI Systems & Robotics Intern — Hybrid, Impactful Projects
AI Systems & Robotics Intern — Hybrid, Impactful Projects

Veeda AI • Zürich

Hybrid
CHF 17,000 - 28,000
Internship
Internship

Veeda AI • Zürich

Hybrid
CHF 17,000 - 28,000