Staff ML Engineer – AWS Trainium & SageMaker

Robots & Pencils

Northern (KY)

Hybrid

USD 140,000 - 230,000

Full time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Robots & Pencils is seeking an engineer to operate and train models on Amazon SageMaker with AWS Trainium. This role involves digging into hardware‑level details, PyTorch training code, and production SageMaker pipelines to ship real solutions for enterprise clients.

You will work hands-on with PyTorch, SageMaker training/inference, and hardware‑specific compilation, collaborating with client teams to scope and deliver production training workloads rather than experiments.

Qualifications

  • Strong, hands‑on PyTorch experience, ideally including distributed or multi‑device training.
  • Production experience with Amazon SageMaker for training and/or inference.
  • Comfort working close to the hardware layer: you understand device‑specific compilation and can debug issues that are actually about the accelerator, not just the model.
  • AWS Trainium or Inferentia (Neurons) experience is a plus.
  • Solid Python fundamentals and comfort operating in a client‑facing, production engineering environment.

Responsibilities

  • Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute.
  • Write and optimize PyTorch training code with a real understanding of how it compiles and executes on Trainium (NeuronCore architecture, compiler behavior, memory and throughput tradeoffs)
  • Diagnose training run issues that show up specifically because of the hardware, not just the model, distinguishing a data or code problem from a compiler or device‑level one
  • Translate a request for "a Trainium job" into an actual working, cost‑aware training pipeline, end to end
  • Tune distributed training runs for throughput and cost on SageMaker's training infrastructure
  • Work directly with client and internal engineering teams to scope and deliver real production training workloads, not experiments that stay in a notebook

Skills

PyTorch
Distributed training
Python

Tools

AWS SageMaker
AWS Trainium
Inferentia

Job description

US Remote

Robots & Pencils is an AWS Partner building production AI systems for enterprise clients who need real engineering, not proofs of concept that never ship. We work forward-deployed, embedded directly with client teams, solving the problems that are too new or too specialized for a typical vendor relationship to handle.

The Role

We're looking for an engineer who can operate and train models on Amazon SageMaker running on AWS Trainium, AWS's custom silicon built specifically for large-scale model training. This isn't a role where you call an API and wait. You'll be walking up the stack: understanding what a training request actually looks like at the Trainium hardware and compiler level, then carrying that understanding all the way up through PyTorch training code and into a production SageMaker pipeline. PyTorch is the backbone of this work. If you know the framework deeply and you're comfortable reasoning about how your code actually behaves on custom accelerator hardware rather than treating it as a black box, this role is built around that skill set specifically.

What You'll Do
  • Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute
  • Write and optimize PyTorch training code with a real understanding of how it compiles and executes on Trainium (NeuronCore architecture, compiler behavior, memory and throughput tradeoffs)
  • Diagnose training run issues that show up specifically because of the hardware, not just the model, distinguishing a data or code problem from a compiler or device-level one
  • Translate a request for "a Trainium job" into an actual working, cost‑aware training pipeline, end to end
  • Tune distributed training runs for throughput and cost on SageMaker's training infrastructure
  • Work directly with client and internal engineering teams to scope and deliver real production training workloads, not experiments that stay in a notebook
What You’re Bring
  • Strong, hands‑on PyTorch experience, ideally including distributed or multi‑device training
  • Production experience with Amazon SageMaker for training and/or inference
  • Comfort working close to the hardware layer: you understand device‑specific compilation and can debug issues that are actually about the accelerator, not just the model
  • AWS Trainium or Inferentia (Ne…
  • Solid Python fundamentals and comfort operating in a client‑facing, production engineering environment
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Engineer – AWS Trainium & SageMaker
Staff ML Engineer – AWS Trainium & SageMaker

Robots & Pencils • United States

On-site
USD 140,000 - 200,000
Staff ML Engineer – AWS Trainium & SageMaker
Staff ML Engineer – AWS Trainium & SageMaker

Robots and Pencils • United States

Remote
USD 140,000 - 200,000
Staff ML Engineer: Production SageMaker & Trainium
Staff ML Engineer: Production SageMaker & Trainium

Robots & Pencils • United States

On-site
USD 140,000 - 200,000
Staff ML Engineer: Production SageMaker on AWS Trainium
Staff ML Engineer: Production SageMaker on AWS Trainium

Robots and Pencils • United States

Remote
USD 140,000 - 200,000
Staff ML Engineer: Production AI on SageMaker w/ Trainium
Staff ML Engineer: Production AI on SageMaker w/ Trainium

Robots & Pencils • Northern (KY)

Hybrid
USD 140,000 - 230,000
AI/ML Inference Engineer for PyTorch on Trainium
AI/ML Inference Engineer for PyTorch on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 150,000 - 210,000
ML Research Engineer, Training
ML Research Engineer, Training

Weave Robotics • San Francisco (CA)

On-site
USD 180,000 - 250,000
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Software Development Engineer II, AWS SageMaker AI
Software Development Engineer II, AWS SageMaker AI

Amazon • Bellevue (WA)

On-site
USD 144,000 - 194,000
RSUs
Health insurance
401(k) matching
Software Engineer - AI/ML, AWS Neuron Apps
Software Engineer - AI/ML, AWS Neuron Apps

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 120,000 - 160,000