ML Systems Engineer — Trainium Inference & Kernels

United States Digital Space LLC

United States

Remote

USD 140,000 - 210,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC seeks a Software Engineer, Trainium, to advance inference workloads on AWS Trainium and build the software stack for frontier models. You will work across kernels, compilers, and model execution to optimize performance on Trainium.

The role involves cross-stack work, identifying bottlenecks and delivering scalable solutions, with a focus on high-performance kernel development and compiler enhancements for accelerator hardware.

Qualifications

  • 3+ years of relevant engineering experience in ML systems, compilers, kernels, runtimes, or performance engineering.
  • Strong systems programming fundamentals and performance-critical software development.
  • Experience with GPU/TPU/Trainium or other accelerator architectures.
  • Ability to reason about performance across multiple layers of the stack, from hardware to ML frameworks.
  • Owning technically ambiguous problems end-to-end and learning new hardware/software domains as needed.
  • Bonus: experience with AWS Trainium or AWS Neuron SDK.
  • Bonus: contributions to ML frameworks such as PyTorch or JAX, compiler infra such as LLVM/MLIR/XLA/Triton.

Responsibilities

  • Build and optimize the company's inference stack for AWS Trainium.
  • Develop high-performance kernels for critical model operations and workloads.
  • Extend and improve compiler support to efficiently target Trainium hardware.
  • Build the systems necessary to execute and optimize the model forward pass on Trainium.
  • Profile workloads and identify bottlenecks across kernels, compiler-generated code, runtime, and model execution.
  • Partner with inference and ML systems teams to bring new models and architectures onto Trainium.
  • Work across the hardware/software boundary to unlock performance and capabilities from specialized AI accelerators.
  • Own complex performance and systems problems end-to-end, from investigation through production deployment.

Skills

ML systems
Kernels
Runtimes
Systems programming
Accelerator architectures
Ambiguity ownership
AWS Trainium/Neuron
PyTorch/JAX/LLVM/MLIR/XLA/Triton

Tools

LLVM
MLIR
XLA
Triton
PyTorch
JAX

Job description

United States Digital Space LLC seeks a Software Engineer, Trainium, to advance inference workloads on AWS Trainium and build the software stack for frontier models. You will work across kernels, compilers, and model execution to optimize performance on Trainium.

The role involves cross-stack work, identifying bottlenecks and delivering scalable solutions, with a focus on high-performance kernel development and compiler enhancements for accelerator hardware.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer: Trainium Inference & Kernels
ML Systems Engineer: Trainium Inference & Kernels

Slope • San Francisco (CA)

On-site
USD 140,000 - 210,000
ML Systems Engineer for Trainium Inference
ML Systems Engineer for Trainium Inference

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
ML Systems Engineer: Optimizing Training & GPU Kernels
ML Systems Engineer: Optimizing Training & GPU Kernels

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
AI/ML Software Engineer — Trainium Distributed Training
AI/ML Software Engineer — Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
AI/ML Performance Engineer – Training on Trainium
AI/ML Performance Engineer – Training on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 140,000 - 210,000
Health Insurance
Medical Insurance
Dental Insurance
+14
AI Infrastructure Kernel Engineer for Large-Scale Training
AI Infrastructure Kernel Engineer for Large-Scale Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Kernel Performance Engineer for Neuron Accelerators
ML Kernel Performance Engineer for Neuron Accelerators

Amazon • Cupertino (CA)

On-site
USD 140,000 - 210,000
Engineering Manager, ML Kernel Performance
Engineering Manager, ML Kernel Performance

Amazon • Cupertino (CA)

On-site
USD 212,700 - 287,700
Health insurance
401(k) matching
Parental leave
+2