ML Systems Engineer: Trainium Inference & Kernels

Slope

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI in San Francisco seeks an experienced Software Engineer to help bring inference workloads to AWS Trainium and build the software stack to run frontier models efficiently on the platform. This deeply technical, cross‑stack role covers kernels, compilers, and model execution.

You will develop high‑performance kernels, improve compiler support, and ensure scalable execution of the model forward pass on Trainium.

Qualifications

  • 3+ years of relevant engineering experience in ML systems, compilers, kernels, runtimes, or performance engineering.
  • Strong systems programming fundamentals and performance‑critical software experience.
  • Experience with GPU, TPU, Trainium, or other accelerator architectures.
  • Ability to reason about performance across stack layers from hardware to ML frameworks.
  • Bonus: experience with AWS Trainium or the AWS Neuron SDK.
  • Bonus: contributions to ML frameworks or compiler infrastructure like LLVM/MLIR/XLA/Triton.

Responsibilities

  • Build and optimize OpenAI's inference stack for AWS Trainium.
  • Develop high‑performance kernels for critical model operations and workloads.
  • Extend and improve compiler support to efficiently target Trainium hardware.
  • Build the systems necessary to execute and optimize the model forward pass on Trainium.
  • Profile workloads and identify bottlenecks across kernels, compiler‑generated code, runtime, and model execution.
  • Partner with inference and ML systems teams to bring new models and architectures onto Trainium.
  • Work across the hardware/software boundary to unlock performance and capabilities from specialized AI accelerators.
  • Own complex performance and systems problems end‑to‑end, from investigation through production deployment.

Skills

ML systems
compilers
kernels
runtimes
performance engineering

Job description

OpenAI in San Francisco seeks an experienced Software Engineer to help bring inference workloads to AWS Trainium and build the software stack to run frontier models efficiently on the platform. This deeply technical, cross‑stack role covers kernels, compilers, and model execution.

You will develop high‑performance kernels, improve compiler support, and ensure scalable execution of the model forward pass on Trainium.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer for Trainium Inference
ML Systems Engineer for Trainium Inference

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
ML Systems Engineer — Trainium Inference & Kernels
ML Systems Engineer — Trainium Inference & Kernels

United States Digital Space LLC • United States

Remote
USD 140,000 - 210,000
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
AI Infrastructure Kernel Engineer for Large-Scale Training
AI Infrastructure Kernel Engineer for Large-Scale Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
AI/ML Software Engineer — Trainium Distributed Training
AI/ML Software Engineer — Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Software Engineer, Trainium
Software Engineer, Trainium

Slope • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Systems Engineer: Optimizing Training & GPU Kernels
ML Systems Engineer: Optimizing Training & GPU Kernels

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000