GPU-Accelerated ML Training Performance Engineer

Blackrock Neurotech

Salt Lake City (UT)

On-site

USD 120,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Blackrock Neurotech, based in Salt Lake City, UT, seeks an ML Training Performance Engineer to optimize large-scale model training across GPUs and clusters. You’ll work hands-on across Python, C++, and GPU kernels while collaborating with researchers and IT teams to scale compute efficiently.

You will own performance baselines, profiling, and tooling, ensuring numerical correctness and scientific intent as models grow in complexity and data. Occasional on-site work is required.

Qualifications

  • Demonstrated experience improving the performance of substantial deep learning training workloads with measurable gains

Responsibilities

  • Own training performance across single-GPU, multi-GPU, and multi-node workloads with reproducible baselines
  • Profile the full training path to identify bottlenecks and impact
  • Optimize tensor layouts, memory, and precision to fit larger models
  • Write and validate custom GPU kernels using CUDA, Triton, or HIP
  • Improve Python scripts, training configs, batching, and optimizers
  • Design distributed training strategies based on model structure and memory limits
  • Collaborate with researchers on hardware-aware changes and convergence impact
  • Coordinate with infra teams on GPU selection, cloud configs, and containers

Skills

Deep learning optimization
GPU kernel development
Profiling & benchmarking
Python & C++ programming
PyTorch knowledge
Distributed training
Linux GPU environments
Big data / HPC concepts

Education

Bachelor’s degree in CS/Engineering or equivalent

Tools

CUDA
HIP/ROCm
Triton
TensorRT

Job description

Blackrock Neurotech, based in Salt Lake City, UT, seeks an ML Training Performance Engineer to optimize large-scale model training across GPUs and clusters. You’ll work hands-on across Python, C++, and GPU kernels while collaborating with researchers and IT teams to scale compute efficiently.

You will own performance baselines, profiling, and tooling, ensuring numerical correctness and scientific intent as models grow in complexity and data. Occasional on-site work is required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Training Performance Engineer - On-Site in Salt Lake City
ML Training Performance Engineer - On-Site in Salt Lake City

Blackrock-Neurotech • Salt Lake City (UT)

On-site
USD 140,000 - 210,000
AI/ML Training Performance Engineer
AI/ML Training Performance Engineer

Blackrock-Neurotech • Salt Lake City (UT)

On-site
USD 140,000 - 210,000
AI/ML Training Performance Engineer
AI/ML Training Performance Engineer

Blackrock Neurotech • Salt Lake City (UT)

On-site
USD 120,000 - 190,000
Remote GPU Performance Engineer for Large-Scale Models
Remote GPU Performance Engineer for Large-Scale Models

United States Digital Space LLC • United States

Remote
USD 140,000 - 210,000
Five weeks paid leave
Comprehensive healthcare (vision +</p>
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Staff ML Systems Engineer - GPU & Performance
Staff ML Systems Engineer - GPU & Performance

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Equity
Performance bonus
Senior GPU-Accelerated ML Systems Engineer
Senior GPU-Accelerated ML Systems Engineer

NVIDIA Corporation • Austin (TX)

On-site
USD 152,000 - 242,000
Staff ML Performance Engineer - GPU Optimization
Staff ML Performance Engineer - GPU Optimization

Google LLC • Sunnyvale (CA)

On-site
USD 186,000 - 228,000
Lead GPU ML Training Performance Engineer
Lead GPU ML Training Performance Engineer

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 120,000 - 160,000
AI Performance Engineer: Scale ML Throughput & Training
AI Performance Engineer: Scale ML Throughput & Training

InvestedintheMission • Sunnyvale (CA)

On-site
USD 180,000 - 260,000