GPU-Driven ML Training Performance Engineer

Socket.dev

Salt Lake City (UT)

On-site

USD 150,000 - 200,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Blackrock Neurotech is seeking an ML Training Performance Engineer to own and optimize the efficiency and scalability of model training on GPU and cloud infra in Salt Lake City. You will work across the training stack with researchers and data/infra teams to drive end-to-end performance gains, maintaining numerical correctness.

You will profile, optimize, and validate training workloads, develop kernels, and tune distributed strategies to fit larger models within available resources.

Qualifications

  • Proven performance improvements for large DL training workloads.
  • Experience with profiling and benchmarking to validate end-to-end gains.
  • Strong Python and C++ programming and debugging skills.
  • Bachelor-level degree or equivalent practical experience.

Responsibilities

  • Own training performance across single/multi-GPU and multi-node workloads.
  • Profile full training paths to identify bottlenecks and quantify impact.
  • Tune tensor layouts, memory, precision, and activation checkpointing.
  • Develop and validate custom GPU kernels with CUDA, Triton, or HIP/ROCm.
  • Improve training scripts, batching, and optimizer execution without changing behavior.
  • Design distributed training strategies based on model structure and interconnects.
  • Collaborate with researchers on hardware-aware changes and convergence.
  • Coordinate data delivery with prefetching and I/O overlap.
  • Work with IT on GPU selection, cloud configs, and capacity planning.

Skills

Python
C++
GPU understanding
CUDA
PyTorch
Distributed training
Profiling & benchmarking
Performance optimization
System debugging
Tensor concepts

Education

Bachelor's degree in CS/CE/EE or related field

Tools

CUDA
HIP/ROCm
Triton

Job description

Blackrock Neurotech is seeking an ML Training Performance Engineer to own and optimize the efficiency and scalability of model training on GPU and cloud infra in Salt Lake City. You will work across the training stack with researchers and data/infra teams to drive end-to-end performance gains, maintaining numerical correctness.

You will profile, optimize, and validate training workloads, develop kernels, and tune distributed strategies to fit larger models within available resources.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Training Performance Engineer
AI/ML Training Performance Engineer

Socket.dev • Salt Lake City (UT)

On-site
USD 150,000 - 200,000
AI/ML Training Performance Engineer
AI/ML Training Performance Engineer

Blackrock Neurotech • Salt Lake City (UT)

On-site
USD 120,000 - 190,000
Remote GPU Performance Engineer for Large-Scale Models
Remote GPU Performance Engineer for Large-Scale Models

United States Digital Space LLC • United States

Remote
USD 140,000 - 210,000
Five weeks paid leave
Comprehensive healthcare (vision +</p>
Lead GPU ML Training Performance Engineer
Lead GPU ML Training Performance Engineer

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 120,000 - 160,000
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
GPU Performance Engineer: Scale ML Inference & Systems
GPU Performance Engineer: Scale ML Inference & Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Systems ML Engineer: Performance, Cloud/Edge, Equity
Systems ML Engineer: Performance, Cloud/Edge, Equity

Breakout Ventures • Cambridge (MA)

On-site
USD 180,000 - 270,000
Equity
Lunch subsidy
Health insurance
+1
AI Performance Engineer: Scale ML Throughput & Training
AI Performance Engineer: Scale ML Throughput & Training

InvestedintheMission • Sunnyvale (CA)

On-site
USD 180,000 - 260,000