Senior GPU Performance Architect: Distributed Training

Jobtailor

Boston (MA)

On-site

USD 180,000 - 240,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a senior capacity optimization lead to manage GPU and large-scale compute workloads across multi-node training environments. You will own the capacity model, including node counts, assumptions, and error bars, and defend capacity decisions with real financial impact.

You will optimize workload mix, preemptible fractions, and checkpointing economics, while improving training throughput and data locality across InfiniBand/RDMA networks and parallel file systems.

Qualifications

  • Five or more years working with GPU and large-scale compute workloads.
  • Real multi-node experience with distributed training performance and interconnects.
  • System-level knowledge of GPU performance, including modern workloads.
  • Deep understanding of AI training and inference performance.
  • Understanding of step time, MFU, compute/communication overlap, and latency floors.

Responsibilities

  • Own the capacity model, including node counts, assumptions and error bars.
  • Defend capacity decisions when real money is committed.
  • Own workload mix, preemptible fraction, and checkpointing economics over time.
  • Explain performance results clearly to researchers in usable terms.
  • Locate actual performance bottlenecks beyond profiler summaries.
  • Measure and optimize parallel file systems, data locality, and interconnect behavior for multi-node GPU training.

Skills

GPU Workloads
Large-Scale Compute
Performance Bottleneck Analysis
AI Training and Inference
Data Locality Optimization
Capacity Modeling
Multi-Node GPU Training
Training Throughput
InfiniBand Networking
RDMA Networking
Parallel File Systems
Data Center Experience
HPC Background
Error Bar Analysis
Batching

Job description

Jobtailor is seeking a senior capacity optimization lead to manage GPU and large-scale compute workloads across multi-node training environments. You will own the capacity model, including node counts, assumptions, and error bars, and defend capacity decisions with real financial impact.

You will optimize workload mix, preemptible fractions, and checkpointing economics, while improving training throughput and data locality across InfiniBand/RDMA networks and parallel file systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Performance Architect & Optimization Lead
Senior GPU Performance Architect & Optimization Lead

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Member of Technical Staff, Performance & Capacity
Member of Technical Staff, Performance & Capacity

Jobtailor • Boston (MA)

On-site
USD 180,000 - 240,000
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Senior Deep Learning Infra Architect for Large-Scale Training
Senior Deep Learning Infra Architect for Large-Scale Training

Jobtailor • California (MO)

On-site
USD 180,000 - 280,000
Senior GPU Capacity & Optimization Architect
Senior GPU Capacity & Optimization Architect

Epoch Biodesign • United States

On-site
USD 160,000 - 195,000
Competitive compensation
Paid time off
Comprehensive health insurance
+3
Senior Architect for Multi-Generation HPC & AI Systems
Senior Architect for Multi-Generation HPC & AI Systems

Jobtailor • California (MO)

On-site
USD 180,000 - 300,000
ML Systems Engineer: Optimizing Training & GPU Kernels
ML Systems Engineer: Optimizing Training & GPU Kernels

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
Senior GPU Capacity & Optimization Architect
Senior GPU Capacity & Optimization Architect

Crusoe • San Francisco (CA)

On-site
USD 160,000 - 195,000
Competitive compensation and equity packages
Comprehensive health, dental & vision insurance
401(k) Retirement plan with company match
+2
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Staff Engineer, AI Compute & Capacity
Staff Engineer, AI Compute & Capacity

Physical Superintelligence • Boston (MA)

Hybrid
USD 180,000 - 280,000