A complete application in a minute — tailored resume and cover letter, ready to send.
Robots & Pencils is seeking an engineer to operate and train models on Amazon SageMaker with AWS Trainium. This role involves digging into hardware‑level details, PyTorch training code, and production SageMaker pipelines to ship real solutions for enterprise clients.
You will work hands-on with PyTorch, SageMaker training/inference, and hardware‑specific compilation, collaborating with client teams to scope and deliver production training workloads rather than experiments.
US Remote
Robots & Pencils is an AWS Partner building production AI systems for enterprise clients who need real engineering, not proofs of concept that never ship. We work forward-deployed, embedded directly with client teams, solving the problems that are too new or too specialized for a typical vendor relationship to handle.
We're looking for an engineer who can operate and train models on Amazon SageMaker running on AWS Trainium, AWS's custom silicon built specifically for large-scale model training. This isn't a role where you call an API and wait. You'll be walking up the stack: understanding what a training request actually looks like at the Trainium hardware and compiler level, then carrying that understanding all the way up through PyTorch training code and into a production SageMaker pipeline. PyTorch is the backbone of this work. If you know the framework deeply and you're comfortable reasoning about how your code actually behaves on custom accelerator hardware rather than treating it as a black box, this role is built around that skill set specifically.