Stand out for this role — generate a tailored resume and cover letter in about a minute.
Together AI in Amsterdam is building and operating one of the world’s largest GPU fleets for frontier model training and inference. You’ll design automation to provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.
You will help create AI infrastructure agents, fleet intelligence platforms, and internal tools to manage tens of thousands of GPUs autonomously, collaborating across hardware, networking, and AI teams.
If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk3+ years building distributed systems, infrastructure platforms, or large-scale backend softwareStrong systems thinking with the ability to understand problems across hardware and softwareA passion for solving complex infrastructure challenges through softwareExperience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologiesStrong software engineering skills in Python, Go, or RustExperience building platforms, automation systems, or developer infrastructureAn automation-first mindset—if a task is repeated, your instinct is to build a system to eliminate itLove building systems that replace repetitive operational workWant to build technology that powers frontier AI modelsEnjoy solving hard problems with no existing playbookThink of infrastructure as a software engineering problemCare deeply about performance, reliability, and scaleYou’ll Thrive Here If You:GPU infrastructure, CUDA, NCCL, NVLink/NVSwitchAI agents and autonomous infrastructure operationsBare-metal provisioning and lifecycle managementDistributed storage systemsInfiniBand or RoCE networkingHardware health monitoring and predictive failure detectionLarge-scale AI training or inference clusters