Infrastructure Operations Engineer

NexGen Cloud

Greater London

Hybrid

GBP 75,000 - 110,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Discretionary bonus
Flexible working
Wellbeing benefits
25 days holiday

Job summary

NexGen Cloud is seeking an Infrastructure Operations Engineer to own and evolve OpenStack and Kubernetes environments for high-performance GPU workloads. You’ll shape platform design, deployment, and operation with a focus on reliability and scale.

Join a fast-moving team, work across Platform, DevOps, AI, Product, and Support, and shape infrastructure decisions while maintaining strong security and incident response capabilities. Hybrid/remote-friendly depending on role and location.

Qualifications

  • Extensive hands-on Linux systems administration skills with deep troubleshooting ability.
  • Direct experience building racks, cabling, and commissioning hardware in data centers.
  • Hands-on experience working in data centers and on-site hardware deployment.
  • Willingness and ability to travel to Quebec sites as required.
  • Solid understanding of networking and storage systems.

Responsibilities

  • Own design, deployment, and operation of OpenStack and Kubernetes environments for GPU workloads.
  • Build and improve infrastructure using infrastructure-as-code and GitOps practices.
  • Optimize GPU workload scheduling and implement monitoring, logging, and alerting for reliability.
  • Lead incident response and drive continuous platform improvements.
  • Maintain security controls across infrastructure and container layers.

Skills

Linux systems
Servers hardware
Data center operations
Networking basics
GitOps / automation

Job description

ABOUT NEXGEN CLOUD:

NexGen Cloud is the company behind Hyperstack, a full-stack AI cloud serving tens of thousands of customers from AI researchers to enterprises running the world's most compute-intensive workloads. We deliver on-demand and private GPU infrastructure to teams who treat performance as a requirement, not a feature.

We're a tight-knit, fast-moving team working at the cutting edge of AI cloud infrastructure. We practice what we preach, equipping our people with AI at every level so we can solve harder problems, ship faster, and keep raising the bar for what enterprise GPU infrastructure looks like.

THE ROLE: Infrastructure Operations Engineer

This role exists because our platform is scaling quickly — and complexity comes with it. As we expand our OpenStack and Kubernetes environments globally, we need engineers who can take real ownership of how the platform is designed, operated, and improved. You'll have direct ownership over business-critical infrastructure that impacts performance, reliability, and customer experience.

This is not a maintenance role. If you like solving hard problems, owning systems end-to-end, and seeing the impact of your work immediately — you'll enjoy this.

WHAT YOU'LL BE DOING:

Rather than a long checklist, here's what success in this role looks like:

  • Own the design, deployment, and operation of OpenStack and Kubernetes environments — ensuring platform performance, scalability, and resilience for GPU workloads
  • Build and improve infrastructure using infrastructure-as-code and GitOps practices, driving automation across provisioning, deployment, and operational workflows
  • Optimise GPU workload scheduling using Kubernetes and NVIDIA tooling, and implement monitoring, logging, and alerting to ensure platform stability
  • Lead incident response and drive continuous improvement of reliability across the platform
  • Maintain strong security controls across infrastructure and container layers — RBAC, network policies, and tenant isolation
  • Work closely with Platform, DevOps, AI, Product, and Support teams to align infrastructure capabilities with customer and platform requirements
ABOUT YOU:

We're more interested in how you think and work than in a perfect CV. You'll likely bring a combination of the following:

Essential

  • Extensive hands-on Linux systems administration skills and knowledge — genuine depth, not surface-level familiarity
  • Strong, proven experience building servers and racks — you've physically assembled, cabled, and commissioned hardware, not just specified or overseen it
  • Direct hands-on experience physically working in data centres — you've personally stacked and racked hardware on-site
  • A willingness and ability to travel to Quebec sites as required
  • A solid understanding of networking and storage systems

Nice to Have

  • Experience installing, racking, and configuring GPU hardware specifically, ideally including NVIDIA platforms
  • Production experience running OpenStack and/or Kubernetes at scale
  • Experience with infrastructure automation, CI/CD, and Git-based workflows
  • Broader exposure to HPC or large-scale compute environments
  • Contributions to open-source projects
WHAT WE OFFER:
  • Competitive salary and annual discretionary bonus scheme
  • Employee wellbeing benefits
  • 25 days of holiday, plus public holidays
  • Flexible working arrangements (remote or hybrid, depending on role and location)
  • Real ownership and autonomy, with the trust to take initiative and experiment
  • The opportunity to make a visible, meaningful impact as we scale
  • Clear career progression and growth opportunities in a fast-growing company
  • A collaborative, international culture built on trust, transparency, and ownership
  • The chance to help shape NexGen Cloud's team, culture, and future alongside ambitious, mission-driven colleagues
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Engineer - Finland
Infrastructure Engineer - Finland

NexGen Cloud • Greater London

Hybrid
GBP 65,000 - 95,000
Competitive salary and discretionary 1
Bonus scheme
25 days holiday
+3
Senior Infrastructure Engineer
Senior Infrastructure Engineer

NexGen Cloud • Greater London

Hybrid
GBP 90,000 - 120,000
Competitive salary
Annual discretionary bonus
25 days holiday
+2
HPC Cluster Architect
HPC Cluster Architect

NexGen Cloud • United Kingdom

Hybrid
GBP 90,000 - 140,000
Competitive salary
25 days holiday
Remote or hybrid
+2
Technical Project Manager
Technical Project Manager

NexGen Cloud • Greater London

On-site
GBP 75,000 - 110,000
Hybrid work
25 days holiday
Discretionary bonus
+2
Infra Operations Engineer – GPU Cloud & OpenStack
Infra Operations Engineer – GPU Cloud & OpenStack

NexGen Cloud • Greater London

Hybrid
GBP 75,000 - 110,000
Discretionary bonus
Flexible working
Wellbeing benefits
+1
Infrastructure Ops Engineer: OpenStack & Kubernetes | Remote
Infrastructure Ops Engineer: OpenStack & Kubernetes | Remote

NexGen Cloud • Greater London

Hybrid
GBP 65,000 - 95,000
Competitive salary and discretionary 1
Bonus scheme
25 days holiday
+3
Senior Infrastructure Engineer (Linux & Networking) - remote, US hours
Senior Infrastructure Engineer (Linux & Networking) - remote, US hours

NexGen Cloud • United Kingdom

Remote
GBP 60,000 - 80,000
100% home-office
Flexible working hours
Collaboration with a diverse team
+2
Business Development Manager
Business Development Manager

NexGen Cloud Ltd • Greater London

Hybrid
GBP 70,000 - 90,000
Competitive salary and annual discretionary bonus
25 days of holiday plus public holidays
Flexible working arrangements
+1
Strategic Sales Lead, AI Natives
Strategic Sales Lead, AI Natives

NexGen Cloud Ltd • Greater London

Hybrid
GBP 65,000 - 90,000
Competitive salary
Annual discretionary bonus
Flexible working arrangements
+1
Supply Chain Manager
Supply Chain Manager

NexGen Cloud • United Kingdom

Hybrid
GBP 70,000 - 110,000
Private Medical Insurance
Enhanced Pension Scheme
Flexible working arrangements
+2