Turn this role into an interview — a resume and cover letter built around what this employer wants.
Lightning AI seeks an experienced Senior Infrastructure Operations Engineer to scale and operate a next‑generation AI infrastructure platform. The role focuses on reliability, automation, and large‑scale GPU environments within a hybrid/UK-wide remote setup or London hub.
You will work with Linux, AWS, Kubernetes, and automation tools to reduce toil and improve platform resilience, partnering with multiple engineering teams across the company.
Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems-designed to take ideas from research to production with less friction.
Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.
We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.
The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:
Lightning AI is seeking an experienced Senior Infrastructure Operations Engineers to help scale and operate our next-generation AI infrastructure platform. Our InfraOps team sits at the center of reliability, automation, and operational scale for GPU infrastructure. This team owns break/fix operations, incident response, customer provisioning, observability, and the automation systems that keep complex infrastructure running efficiently. This role may be fully remote within the UK, or hybrid out of one of our London office hub, with occasional team and company offsites. We are not able to provide work visa sponsorship for this position at this time.
In this role, you'll work hands-on with large-scale GPU environments, Linux systems, bare metal infrastructure, provisioning workflows, and platform reliability. You'll partner closely with Infrastructure Engineering, Network Operations, and Software Platform teams to troubleshoot issues, improve operational efficiency, and build automation that reduces manual toil over time.