ML Systems Engineer for Trainium Inference

AI Chopping Block

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 280,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking a Software Engineer, Trainium, to bring OpenAI's inference workloads to AWS Trainium and build the software stack for efficient model execution.

This deeply technical role spans kernels, compilers, and model execution, addressing performance bottlenecks and enabling frontier-scale AI systems across the stack.

If you enjoy ML systems, compilers, and accelerator hardware challenges, this role fits you—self-directed and capable across abstraction layers.

Qualifications

  • 3+ years of relevant engineering experience in ML systems, compilers, kernels, runtimes, or performance engineering.
  • Experience with GPU/TPU/Trainium or other specialized accelerators.
  • Ability to reason about performance across hardware, kernels, compilers and ML frameworks.
  • Owning ambiguous problems end-to-end and learning new domains as needed.
  • Bonus: experience with AWS Trainium or AWS Neuron SDK; contributions to ML frameworks or compiler infra.

Responsibilities

  • Build and optimize OpenAI's inference stack for AWS Trainium.
  • Develop high-performance kernels for critical model operations.
  • Extend and improve compiler support to target Trainium hardware.
  • Build systems to execute and optimize the model forward pass on Trainium.
  • Profile workloads and identify bottlenecks across kernels, runtime, and model execution.
  • Collaborate with inference and ML systems teams to bring models and architectures to Trainium.
  • Work across hardware/software boundaries to unlock accelerator performance.

Skills

Systems programming
Performance engineering
ML systems
Kernel development
Compiler support

Tools

LLVM
MLIR
XLA
Triton
PyTorch

Job description

OpenAI is seeking a Software Engineer, Trainium, to bring OpenAI's inference workloads to AWS Trainium and build the software stack for efficient model execution.

This deeply technical role spans kernels, compilers, and model execution, addressing performance bottlenecks and enabling frontier-scale AI systems across the stack.

If you enjoy ML systems, compilers, and accelerator hardware challenges, this role fits you—self-directed and capable across abstraction layers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer: Trainium Inference & Kernels
ML Systems Engineer: Trainium Inference & Kernels

Slope • San Francisco (CA)

On-site
USD 140,000 - 210,000
ML Systems Engineer — Trainium Inference & Kernels
ML Systems Engineer — Trainium Inference & Kernels

United States Digital Space LLC • United States

Remote
USD 140,000 - 210,000
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
Software Engineer, Trainium
Software Engineer, Trainium

Slope • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI/ML Software Engineer — Trainium Distributed Training
AI/ML Software Engineer — Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
AI/ML Performance Engineer – Training on Trainium
AI/ML Performance Engineer – Training on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 140,000 - 210,000
Health Insurance
Medical Insurance
Dental Insurance
+14
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Software Engineer, Trainium
Software Engineer, Trainium

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
GenAI ML Systems Engineer
GenAI ML Systems Engineer

Amazon Web Services (AWS) • New York (NY)

On-site
USD 158,000 - 214,000
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000