Staff ML Engineer: Production SageMaker on AWS Trainium

Robots and Pencils

United States

Remote

USD 140,000 - 200,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Robots & Pencils is seeking an engineer who can operate and train models on Amazon SageMaker with AWS Trainium, pushing beyond API calls into hardware-aware training.

You'll work across Torch, compiler behavior, and production SageMaker pipelines, crafting end-to-end training solutions for enterprise clients.

If you love PyTorch, hardware details, and building scalable ML systems, join our forward-deployed team delivering real production workloads.

Qualifications

  • Strong, hands-on PyTorch experience, ideally including distributed or multi-device training.
  • Production experience with Amazon SageMaker for training and/or inference.
  • Comfort working close to the hardware layer: you understand device-specific compilation and can debug issues that are actually about the accelerator, not just the model.
  • AWS Trainium or Inferentia (Neuron SDK) experience is a strong plus; if you don’t have it yet but have deep PyTorch and a track record of picking up new hardware targets fast, we want to talk to you
  • Solid Python fundamentals and comfort operating in a client-facing, production engineering environment.

Responsibilities

  • Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute
  • Write and optimize PyTorch training code with a real understanding of how it compiles and executes on Trainium (NeuronCore architecture, compiler behavior, memory and throughput tradeoffs)
  • Diagnose training run issues that show up specifically because of the hardware, not just the model, distinguishing a data or code problem from a compiler or device-level one
  • Translate a request for “a Trainium job” into an actual working, cost-aware training pipeline, end to end
  • Tune distributed training runs for throughput and cost on SageMaker’s training infrastructure
  • Work directly with client and internal engineering teams to scope and deliver real production training workloads, not experiments that stay in a notebook

Skills

PyTorch
SageMaker
Python
Distributed training
Hardware acceleration

Tools

SageMaker
Trainium
Neuron SDK
PyTorch

Job description

Robots & Pencils is seeking an engineer who can operate and train models on Amazon SageMaker with AWS Trainium, pushing beyond API calls into hardware-aware training.

You'll work across Torch, compiler behavior, and production SageMaker pipelines, crafting end-to-end training solutions for enterprise clients.

If you love PyTorch, hardware details, and building scalable ML systems, join our forward-deployed team delivering real production workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Engineer: Production SageMaker & Trainium
Staff ML Engineer: Production SageMaker & Trainium

Robots & Pencils • United States

On-site
USD 140,000 - 200,000
Staff ML Engineer: Production AI on SageMaker w/ Trainium
Staff ML Engineer: Production AI on SageMaker w/ Trainium

Robots & Pencils • Northern (KY)

Hybrid
USD 140,000 - 230,000
Staff ML Engineer – AWS Trainium & SageMaker
Staff ML Engineer – AWS Trainium & SageMaker

Robots and Pencils • United States

Remote
USD 140,000 - 200,000
Staff ML Engineer – AWS Trainium & SageMaker
Staff ML Engineer – AWS Trainium & SageMaker

Robots & Pencils • United States

On-site
USD 140,000 - 200,000
Staff ML Engineer – AWS Trainium & SageMaker
Staff ML Engineer – AWS Trainium & SageMaker

Robots & Pencils • Northern (KY)

Hybrid
USD 140,000 - 230,000
Senior SDE – SageMaker Training: Scale ML Customization
Senior SDE – SageMaker Training: Scale ML Customization

Amazon Inc. • Bellevue (WA)

On-site
USD 168,000 - 227,000
Health insurance
RSUs
Paid time off
+1
Lead ML Systems Engineer – AI/GenAI Acceleration
Lead ML Systems Engineer – AI/GenAI Acceleration

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs (restricted stock units)
401(k) matching
+2
AI/ML Inference Engineer for PyTorch on Trainium
AI/ML Inference Engineer for PyTorch on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 150,000 - 210,000
Senior ML Engineer — Build Production AI (Remote)
Senior ML Engineer — Build Production AI (Remote)

Tekmetric • United States

Hybrid
USD 140,000 - 190,000
Remote work flexibility
Competitive base salary
Generous PTO
+3
Senior ML/AI Engineer — End-to-End Production ML
Senior ML/AI Engineer — End-to-End Production ML

Toyota • Plano (AL)

On-site
USD 180,000 - 240,000
Healthcare plans
Tuition reimbursement
Vehicle Purchase Discount
+1