ML Training Performance Engineer - On-Site in Salt Lake City

Blackrock-Neurotech

Salt Lake City (UT)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Blackrock Neurotech is seeking an ML Training Performance Engineer to own the efficiency and scalability of neural network training on GPU and cloud infrastructure. You will work as an individual contributor on a small research team across the training stack, partnering with model researchers, data engineers, and IT to improve performance while preserving scientific integrity.

The role is on-site at our Salt Lake City headquarters.

Qualifications

  • Bachelor’s degree or equivalent practical experience with 3+ years in industry
  • Demonstrated performance improvements in large DL training workloads
  • Measurement-driven approach with profiling/benchmarking
  • Strong Python and C++ programming and debugging
  • Deep understanding of GPU execution and memory hierarchies
  • Experience with CUDA or HIP/ROCm kernels
  • Proficiency in PyTorch and mixed-precision training
  • Strong understanding of DL architectures and training computations
  • Experience with distributed training and GPU interconnects
  • Experience configuring Linux-based GPU environments and containers

Responsibilities

  • Own training performance across single-GPU, multi-GPU, and multi-node workloads with reproducible baselines
  • Profile full training path to identify bottlenecks and prioritize improvements
  • Optimize tensor layouts, memory, and execution graphs for larger models
  • Develop and validate CUDA/HIP kernels when needed for performance
  • Improve Python training scripts, framework settings, batching, and optimizers
  • Design and tune distributed training strategies based on model structure and memory
  • Collaborate with researchers on hardware-aware architecture and hyperparameters
  • Coordinate with infrastructure engineers on prefetching and data I/O overlap
  • Work with IT on GPUs, cloud instances, networking, and containers
  • Build robust checkpoint/restart/recovery workflows and regression checks
  • Communicate benchmark results and resource recommendations clearly

Skills

Python
C++
GPU perf optimization
PyTorch
Distributed training
Profiling
Linux
Collaboration
CUDA

Education

Bachelor’s degree in CS/CE or related field
Master’s degree in related field
PhD in related field

Tools

CUDA
HIP/ROCm
Triton
Gloo/NCCL
Docker

Job description

Blackrock Neurotech is seeking an ML Training Performance Engineer to own the efficiency and scalability of neural network training on GPU and cloud infrastructure. You will work as an individual contributor on a small research team across the training stack, partnering with model researchers, data engineers, and IT to improve performance while preserving scientific integrity.

The role is on-site at our Salt Lake City headquarters.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU-Accelerated ML Training Performance Engineer
GPU-Accelerated ML Training Performance Engineer

Blackrock Neurotech • Salt Lake City (UT)

On-site
USD 120,000 - 190,000
AI/ML Training Performance Engineer
AI/ML Training Performance Engineer

Blackrock-Neurotech • Salt Lake City (UT)

On-site
USD 140,000 - 210,000
AI/ML Training Performance Engineer
AI/ML Training Performance Engineer

Blackrock Neurotech • Salt Lake City (UT)

On-site
USD 120,000 - 190,000
Lead Neural Data Infrastructure Engineer - BCI
Lead Neural Data Infrastructure Engineer - BCI

Blackrock Neurotech • Salt Lake City (UT)

On-site
USD 120,000 - 180,000
Senior ML Platform Engineer – Healthcare AI & Production
Senior ML Platform Engineer – Healthcare AI & Production

University of Utah Health • Salt Lake City (UT)

On-site
USD 140,000 - 180,000
Health Coverage
Dental Coverage
Life Insurance
+3
Neural Data Platform Engineer — Scale Brain Interfaces
Neural Data Platform Engineer — Scale Brain Interfaces

Blackrock-Neurotech • Salt Lake City (UT)

On-site
USD 130,000 - 190,000
Remote GPU Performance Engineer for Large-Scale Models
Remote GPU Performance Engineer for Large-Scale Models

United States Digital Space LLC • United States

Remote
USD 140,000 - 210,000
Five weeks paid leave
Comprehensive healthcare (vision +</p>
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2
Systems ML Engineer: Performance, Cloud/Edge, Equity
Systems ML Engineer: Performance, Cloud/Edge, Equity

Breakout Ventures • Cambridge (MA)

On-site
USD 180,000 - 270,000
Equity
Lunch subsidy
Health insurance
+1
Staff ML Performance Engineer: Scale Training & GPU
Staff ML Performance Engineer: Scale Training & GPU

OpenDigital Limited • Sunnyvale (CA), Northern (KY)

Hybrid
USD 336,000 - 359,000