Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.
Lightning AI is seeking an Infrastructure Operations Engineer to scale and operate a cutting-edge AI infrastructure platform supporting large-scale training and inference workloads. The role focuses on GPU infrastructure reliability, automation, and platform performance, spanning Linux servers, bare metal and virtualized environments, and provisioning workflows.
You will own on-call rotations, deploy updates, and collaborate with cross-functional teams to reduce toil, improve observability, and
An Infrastructure Operations Engineer is needed to help scale and operate a next-generation AI infrastructure platform that supports training and inference workloads at large scale. The InfraOps team sits at the center of reliability, automation, and operational scale for GPU infrastructure, owning break/fix operations, incident response, customer provisioning, observability, and the automation systems that keep complex infrastructure running efficiently. The role is hands-on, working across large-scale GPU environments, Linux systems, bare metal infrastructure, provisioning workflows, and platform reliability.