Get more replies from employers
Send a job-specific resume in minutes.
Watney is seeking an ML Infrastructure Engineer to turn live fleet data into trainable models. You will own training and inference infrastructure, build data pipelines, and optimize GPU utilization for large-scale training runs.
The role requires strong Python and PyTorch or TensorFlow experience and a track record of scalable ML systems. Join a growing robotics company applying autonomous systems to expand human capability in data center construction.
Expand human ambition in the physical world.
Critical infrastructure is constrained by labor shortages, hazardous working conditions, and operational complexity. Watney builds and deploys autonomous robotic systems that increase the speed and capacity of buildout, starting with data centers.
At Watney, ML Infrastructure engineers turn data collected from a live fleet of robots into better models. The fleet produces large volumes of video and telemetry data from real work in the field, and making that data trainable is one of the hardest systems problems at the company.
As we continue to scale, these systems will require larger training runs with more data, expanded clusters, and optimal GPU utilization.
Own training and inference infrastructure
Build the data pipelines that these training runs depend on
Make experiments fast to launch and reproduce
Contribute to our core training code
Have built ML infrastructure that carried real production training runs
Have scaled distributed training systems
Strong experience with Python, PyTorch or TensorFlow
Have experience identifying and troubleshooting GPU performance bottlenecks in large-scale training environments
We’re committed to building a diverse, inclusive team. At Watney Robotics, we welcome people of all backgrounds and identities, and we make hiring decisions based on skills, experience, and potential. If you’re passionate about robotics but don’t meet every requirement, we still encourage you to apply!
Follow us here onXandLinkedIn