An application made for this job — a tailored resume and cover letter that speak straight to the posting.
NexGen Cloud in London is seeking an Infrastructure Operations Engineer to design, deploy and operate OpenStack and Kubernetes environments for GPU workloads, ensuring performance, scalability and reliability. You will implement IaC and GitOps, drive automation, lead incident response, maintain security controls, and collaborate across Platform, DevOps, AI, Product and Support teams.
Some travel to Quebec sites may be required.
NexGen Cloud is the company behind Hyperstack, a full-stack AI cloud serving tens of thousands of customers from AI researchers to enterprises running the world's most compute-intensive workloads. We deliver on-demand and private GPU infrastructure to teams who treat performance as a requirement, not a feature.
We're a tight-knit, fast-moving team working at the cutting edge of AI cloud infrastructure. We practice what we preach, equipping our people with AI at every level so we can solve harder problems, ship faster, and keep raising the bar for what enterprise GPU infrastructure looks like.
This role exists because our platform is scaling quickly — and complexity comes with it. As we expand our OpenStack and Kubernetes environments globally, we need engineers who can take real ownership of how the platform is designed, operated, and improved. You'll have direct ownership over business-critical infrastructure that impacts performance, reliability, and customer experience.
This is not a maintenance role. If you like solving hard problems, owning systems end-to-end, and seeing the impact of your work immediately — you'll enjoy this.
Rather than a long checklist, here's what success in this role looks like:
We're more interested in how you think and work than in a perfect CV. You'll likely bring a combination of the following:
Essential
Nice to Have