AI/ML Training Performance Engineer

Blackrock-Neurotech

Salt Lake City (UT)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Blackrock Neurotech is seeking an ML Training Performance Engineer to own the efficiency and scalability of neural network training on GPU and cloud infrastructure. You will work as an individual contributor on a small research team across the training stack, partnering with model researchers, data engineers, and IT to improve performance while preserving scientific integrity.

The role is on-site at our Salt Lake City headquarters.

Qualifications

  • Bachelor’s degree or equivalent practical experience with 3+ years in industry
  • Demonstrated performance improvements in large DL training workloads
  • Measurement-driven approach with profiling/benchmarking
  • Strong Python and C++ programming and debugging
  • Deep understanding of GPU execution and memory hierarchies
  • Experience with CUDA or HIP/ROCm kernels
  • Proficiency in PyTorch and mixed-precision training
  • Strong understanding of DL architectures and training computations
  • Experience with distributed training and GPU interconnects
  • Experience configuring Linux-based GPU environments and containers

Responsibilities

  • Own training performance across single-GPU, multi-GPU, and multi-node workloads with reproducible baselines
  • Profile full training path to identify bottlenecks and prioritize improvements
  • Optimize tensor layouts, memory, and execution graphs for larger models
  • Develop and validate CUDA/HIP kernels when needed for performance
  • Improve Python training scripts, framework settings, batching, and optimizers
  • Design and tune distributed training strategies based on model structure and memory
  • Collaborate with researchers on hardware-aware architecture and hyperparameters
  • Coordinate with infrastructure engineers on prefetching and data I/O overlap
  • Work with IT on GPUs, cloud instances, networking, and containers
  • Build robust checkpoint/restart/recovery workflows and regression checks
  • Communicate benchmark results and resource recommendations clearly

Skills

Python
C++
GPU perf optimization
PyTorch
Distributed training
Profiling
Linux
Collaboration
CUDA

Education

Bachelor’s degree in CS/CE or related field
Master’s degree in related field
PhD in related field

Tools

CUDA
HIP/ROCm
Triton
Gloo/NCCL
Docker

Job description

Build the systems that expand human capability

At Blackrock Neurotech, we’ve spent decades making the impossible possible – helping people move, speak, and reconnect with the world when they otherwise could not. We’ve seen that restoring function restores more than ability. It restores independence, identity, and agency.

Today, we are building the next generation of human capability: brain-computer interfaces that are designed to be safe, scalable, and trusted in the real world. Our work is not only about reconnecting people to what was lost, but about expanding what is possible – creating a seamless interface between human intent and technology.

This is foundational work in a category-defining field. You will help build the infrastructure for a future where neural interfaces are invisible, reliable, and deeply human-centered.

Working at Blackrock Neurotech means:
  • Owning meaningful, high-impact problems at the frontier of science and engineering
  • Building alongside experienced, thoughtful peers across disciplines
  • Solving technically complex challenges grounded in real human outcomes
  • Contributing to a culture that values rigor, clarity, and long-term thinking over noise
The Role

The ML Training Performance Engineer will own the efficiency and scalability of training models on GPU and cloud infrastructure. You will turn available compute into faster, more capable experiments as our model training scales in complexity and compute requirements.

As a hands‑on individual contributor on a small research team, you will work across the training stack, from Python and model execution to GPU kernels, distributed communication, and runtime environments. Partnering with model researchers, data engineers, and infrastructure and IT teams, you will identify and implement performance improvements while preserving numerical correctness and scientific intent.

You will have significant ownership over how we measure, optimize, and scale training performance, establishing the baselines, tooling, and technical approaches that will support our AI/ML work as it grows.

What You’ll Do
  • Own training performance across single‑GPU, multi‑GPU, and multi‑node workloads, establishing reproducible baselines for throughput, memory use, utilization, time to target quality, and cost
  • Profile the full training path to distinguish compute, memory, communication, CPU, and I/O bottlenecks and prioritize changes with measurable end‑to‑end impact
  • Optimize tensor layouts, precision, memory allocation, activation checkpointing, operator fusion, and execution graphs to fit larger or longer‑context models within available resources
  • Write, tune, and validate custom GPU kernels using CUDA, Triton, HIP/ROCm, or the appropriate platform tools when existing implementations limit performance
  • Improve Python training scripts, framework and compiler settings, batching, gradient accumulation, and optimizer execution while preserving intended training behavior
  • Design and tune distributed training strategies, including data, tensor, pipeline, or sharded parallelism, based on model structure, memory limits, and interconnect topology
  • Partner with model researchers on hardware‑aware architecture and hyperparameter changes, measuring their effects on convergence, model quality, and compute requirements
  • Coordinate with the neural data infrastructure engineer on prefetching, pinned memory, host‑to‑device transfer, and I/O overlap so data delivery keeps pace with training
  • Work with infrastructure and IT on GPU selection, cloud instance configurations, networking, drivers, containers, scheduling, and capacity planning as training needs and compute capacity scale
  • Build robust checkpoint, restart, and recovery workflows and performance regression checks that keep long‑running experiments reproducible and productive
  • Communicate benchmark evidence, numerical tradeoffs, scaling limits, and resource recommendations clearly to researchers and organizational stakeholders
What You Bring
  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent practical experience with 3+ years of relevant industry experience OR Master’s degree with 1+ years of relevant industry experience OR PhD in a related field
  • Demonstrated experience improving the performance of substantial deep learning training workloads, with measured gains in speed, memory efficiency, or compute costI
  • A measurement‑driven approach to performance optimization, using profiling and benchmarking to validate meaningful end‑to‑end improvements
  • Exceptional programming ability in Python and C++ or a comparable systems language, with strong debugging, testing, and performance analysis practices
  • Deep understanding of GPU execution, including memory hierarchies, memory coalescing, thread blocks, warps or wavefronts, occupancy, synchronization, and bandwidth limits
  • Hands‑on experience developing and profiling GPU kernels with CUDA or HIP/ROCm, and the ability to diagnose correctness and performance at the hardware level
  • Deep knowledge of PyTorch or an equivalent framework, including automatic differentiation, computation graphs, tensor storage, compilation, and mixed‑precision training
  • Strong understanding of deep learning architectures and the underlying computations that drive training performance
  • Experience with distributed training, collective communication, sharding, and the interaction between model partitioning and GPU interconnects
  • Strong understanding of numerical stability and the ability to validate gradients, convergence, and model quality after performance changes
  • Experience configuring and diagnosing Linux‑based GPU environments, containers, cloud compute, and high‑throughput storage or networking
  • Ability to collaborate closely with researchers and infrastructure teams and make clear tradeoffs between implementation effort, performance reliability, and scientific value
  • Experience with Triton, compiler optimization, advanced GPU profiling tools, multiple accelerator generations, long‑sequence or multimodal models, neural time series, or large‑scale model training is a plus

Working Location:

This is an on‑site role based at Blackrock Neurotech's headquarters in Salt Lake City, Utah. Occasional travel may be required.

How We Work

We are a small, experienced team working on consequential problems.

  • We take ownership of outcomes and follow through with clarity and accountability
  • We prioritize sustained, high‑quality work over performative urgency
  • We value rigor, sound judgement and thoughtful decision‑making
  • We collaborate deliberately: low ego, high trust and high context

This is a high‑ownership role, but it is not an "always-on" one. We expect strong work and our people to have a life outside of it.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI/ML Training Performance Engineer
AI/ML Training Performance Engineer

Blackrock Neurotech • Salt Lake City (UT)

On-site
USD 120,000 - 190,000
Neural Data Infrastructure Engineer
Neural Data Infrastructure Engineer

Blackrock-Neurotech • Salt Lake City (UT)

On-site
USD 130,000 - 190,000
Neural Data Infrastructure Engineer
Neural Data Infrastructure Engineer

Blackrock Neurotech • Salt Lake City (UT)

On-site
USD 120,000 - 180,000
GPU-Accelerated ML Training Performance Engineer
GPU-Accelerated ML Training Performance Engineer

Blackrock Neurotech • Salt Lake City (UT)

On-site
USD 120,000 - 190,000
ML Training Performance Engineer - On-Site in Salt Lake City
ML Training Performance Engineer - On-Site in Salt Lake City

Blackrock-Neurotech • Salt Lake City (UT)

On-site
USD 140,000 - 210,000
AI Performance Engineer
AI Performance Engineer

applied • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Founding Senior AI Infrastructure Engineer
Founding Senior AI Infrastructure Engineer

Goaly • Palo Alto (CA)

On-site
USD 140,000 - 210,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Training Performance Engineer
Training Performance Engineer

Slope • San Francisco (CA)

On-site
USD 250,000 - 460,000
Relocation assistance
Flexible working hours
Collaborative work environment
Member of Technical Staff - AI Training Platform
Member of Technical Staff - AI Training Platform

Unconventional AI • Los Angeles (CA), California (MO)

On-site
USD 180,000 - 260,000
Competitive salary and equity
Best-in-class health benefits
401k matching
+2