Staff ML Engineer: Production SageMaker & Trainium

Robots & Pencils

United States

On-site

USD 140,000 - 200,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Robots & Pencils is seeking an engineer who can operate and train models on Amazon SageMaker with AWS Trainium, integrating from hardware-level concepts through PyTorch to production pipelines. You’ll work hands-on beyond APIs, shaping training requests at the accelerator and compiler level.

Strong PyTorch expertise, production experience with SageMaker, and comfort near the hardware layer are required. You’ll diagnose issues, optimize distributed training, and collaborate with client teams to

Qualifications

  • Hands-on PyTorch experience with distributed or multi-device training.
  • Production experience with Amazon SageMaker for training and inference.
  • Comfort working close to accelerator hardware: understanding compilation and device behavior.

Responsibilities

  • Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute.
  • Write and optimize PyTorch training code with a real understanding of how it compiles and executes on Trainium (NeuronCore architecture, compiler behavior, memory and throughput tradeoffs).
  • Diagnose training run issues that arise from hardware, not just the model, distinguishing data or code problems from compiler or device-level ones.
  • Translate a request for "a Trainium job" into an actual working, cost-aware training pipeline, end to end.
  • Tune distributed training runs for throughput and cost on SageMaker's training infrastructure.
  • Work directly with client and internal engineering teams to scope and deliver real production training workloads.

Skills

PyTorch
Python
Distributed training
SageMaker

Tools

AWS Trainium
Neuron SDK
SageMaker

Job description

Robots & Pencils is seeking an engineer who can operate and train models on Amazon SageMaker with AWS Trainium, integrating from hardware-level concepts through PyTorch to production pipelines. You’ll work hands-on beyond APIs, shaping training requests at the accelerator and compiler level.

Strong PyTorch expertise, production experience with SageMaker, and comfort near the hardware layer are required. You’ll diagnose issues, optimize distributed training, and collaborate with client teams to

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Engineer: Production SageMaker on AWS Trainium
Staff ML Engineer: Production SageMaker on AWS Trainium

Robots and Pencils • United States

Remote
USD 140,000 - 200,000
Staff ML Engineer: Production AI on SageMaker w/ Trainium
Staff ML Engineer: Production AI on SageMaker w/ Trainium

Robots & Pencils • Northern (KY)

Hybrid
USD 140,000 - 230,000
Staff ML Engineer – AWS Trainium & SageMaker
Staff ML Engineer – AWS Trainium & SageMaker

Robots and Pencils • United States

Remote
USD 140,000 - 200,000
Staff ML Engineer – AWS Trainium & SageMaker
Staff ML Engineer – AWS Trainium & SageMaker

Robots & Pencils • Northern (KY)

Hybrid
USD 140,000 - 230,000
Staff ML Engineer – AWS Trainium & SageMaker
Staff ML Engineer – AWS Trainium & SageMaker

Robots & Pencils • United States

On-site
USD 140,000 - 200,000
Senior SDE – SageMaker Training: Scale ML Customization
Senior SDE – SageMaker Training: Scale ML Customization

Amazon Inc. • Bellevue (WA)

On-site
USD 168,000 - 227,000
Health insurance
RSUs
Paid time off
+1
Lead ML Systems Engineer – AI/GenAI Acceleration
Lead ML Systems Engineer – AI/GenAI Acceleration

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs (restricted stock units)
401(k) matching
+2
Senior ML/AI Engineer — End-to-End Production ML
Senior ML/AI Engineer — End-to-End Production ML

Toyota • Plano (AL)

On-site
USD 180,000 - 240,000
Healthcare plans
Tuition reimbursement
Vehicle Purchase Discount
+1
Senior AI/ML Engineer — Production-Scale Models
Senior AI/ML Engineer — Production-Scale Models

Veriipro • Alpharetta (GA)

On-site
Confidential
Gen AI & ML Solutions Architect — SageMaker Expert
Gen AI & ML Solutions Architect — SageMaker Expert

Amazon Web Services (AWS) • Chicago (IL)

On-site
USD 154,000 - 208,000