Infrastructure Operations Manager (London)

Nscale

Greater London

On-site

GBP 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nscale is seeking an experienced Infrastructure Operations Manager to lead the AI data center’s devices and infrastructure, ensuring operational excellence and on‑time delivery of client SLAs. You will oversee teams, coordinate with vendors, and drive reliability across 24/7 operations.

Ideal candidates have 5+ years in data center management, with HPC/GPU expertise, strong communication skills, and a hands‑on approach to hardware and environmental systems.

Qualifications

  • Bachelor’s degree in CS or Engineering; strong technical foundation.
  • 5+ years managing data centers, especially HPC and GPU environments.
  • Proven leadership in building and developing teams.
  • Technical expertise in HPC, GPU deployments, and related hardware/software.
  • Deep knowledge of power, cooling and environmental systems.

Responsibilities

  • Own site and infrastructure management for the AI data center, ensuring 24/7 reliability.
  • Lead and mentor engineers, plan shifts and on-call coverage.
  • Serve as main client and vendor liaison, report SLA/KPI results.
  • Oversee spare parts inventory and procurement to prevent downtime.
  • Maintain HPC/GPU installations and optimize performance.
  • Monitor operations and propose improvements for reliability and cost savings.

Skills

Leadership
Client-facing communication
HPC deployment
GPU deployments
Power, cooling, environmental systems
Inventory & spare parts management

Education

Bachelor’s degree in Computer Science or Engineering

Tools

NVIDIA GPUs
CUDA

Job description

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

At Nscale, our Operations team plays a critical role in maintaining service availability, driving service reliability and rapid response to customer tickets

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

About The Role

We are looking for an experienced Team Lead to take overall responsibility for the devices and infrastructure in our AI Data Center. The Infrastructure Operations Manager will ensure operational excellence, oversee personnel, maintain critical infrastructure, and meet client-driven SLA and KPI requirements. Reporting to the Head of Infrastructure Operations, this role involves close collaboration with internal teams, clients, and external vendors. It also includes on-call duties and occasional travel to other sites as business needs arise.

What You’ll Be Doing
Site & Infrastructure Management
  • Overall accountability for all devices, systems, and infrastructure at the data center.
  • Ensure 24/7 operational reliability and optimal performance for high-performance AI workloads.
  • Monitor and manage power, cooling, and environmental conditions, addressing any operational risks proactively.
  • Oversee installation, configuration, and maintenance of HPC and GPU systems.
Team Leadership & Development
  • Lead and mentor a team of engineers, technicians, and support staff.
  • Provide training and guidance to junior staff, ensuring the team is equipped to manage site operations effectively.
  • Plan and manage shift schedules to guarantee continuous coverage and on-call availability.
Client & Vendor Relations
  • Serve as the primary POC for clients, providing regular reporting on site SLA and KPI’s
  • Collaborate with vendors and contractors to manage procurement, repairs, and upgrades.
  • Build and maintain strong relationships to ensure smooth operation and project delivery.
Inventory & Resource Management
  • Manage spare parts inventory, ensuring sufficient stock levels to minimize downtime during repairs or upgrades.
  • Track usage and coordinate procurement to avoid supply shortages.
Technical Expertise
  • Maintain a strong working knowledge of HPC and GPU installations, including hardware configuration, network architecture, and performance optimization.
  • Troubleshoot and resolve technical issues, escalating complex problems when necessary.
  • Stay updated on emerging trends and technologies in AI and HPC infrastructure.
Operational Monitoring & Reporting
  • Implement and manage tools to monitor data center performance and resource utilization.
  • Prepare detailed reports for senior management and clients on performance metrics, downtime, and operational efficiency.
  • Propose and implement improvements to enhance reliability, scalability, and cost-effectiveness.
About You
  • Bachelor’s degree in Computer Science, Engineering, or a related field.
  • 5+ years of experience managing data centers, particularly in HPC and GPU environments.
  • Proven leadership skills with experience managing and developing teams.
  • Technical expertise in HPC, GPU deployment, and associated hardware/software solutions.
  • Strong understanding of power, cooling, and environmental systems in data centers.
  • Excellent client-facing communication and reporting skills.
  • Experience with inventory and spare parts management.
Nice To Have
  • Certifications in data center management (e.g., CDCP, CDCS, or similar).
  • Hands-on experience with NVIDIA GPUs, CUDA, and AI frameworks.
  • Familiarity with hybrid cloud/HPC environments.
Work Environment & Requirements
  • Based on-site at the data center with occasional travel to other locations as required.
  • Participate in an on-call rotation to respond to urgent operational issues.
  • Commitment to continuous learning to stay ahead of evolving technologies and standards.
What We Can Offer You

You’ll have the opportunity to help shape the operating standards behind a next-generation AI cloud platform, working on complex infrastructure challenges with real ownership and impact. This is a chance to play a meaningful role in scaling high-performance, sustainable data centre operations in a fast-moving environment.

Equal Opportunities Statement

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enriches our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Deployment Program Manager
Deployment Program Manager

Nscale • Greater London

On-site
GBP 100,000 - 130,000
Senior Manager, Finance Operations - Data Centres
Senior Manager, Finance Operations - Data Centres

Nscale • Greater London

Hybrid
GBP 110,000 - 160,000
Base + equity
Growth plan
Flexible work
Principal Network Engineer
Principal Network Engineer

Nscale • Greater London

On-site
GBP 120,000 - 170,000
Base + equity
Flexible working
Competitive package
+1
Director, Workplace Experience & Real Estate
Director, Workplace Experience & Real Estate

Nscale • City Of London

On-site
GBP 150,000 - 230,000
Competitive compensation package (base + equity)
Clear progression plan
Flexibility in work arrangements
Reliability Engineer SME (M&E)
Reliability Engineer SME (M&E)

Nscale • Greater London

On-site
GBP 90,000 - 120,000
Equity
Bonus
Flexible working
+1
Talent Acquisition Partner, Data Centers EMEA
Talent Acquisition Partner, Data Centers EMEA

Nscale • Greater London

On-site
GBP 70,000 - 120,000
Competitive compensation
Comprehensive benefits
Equity
+2
Senior Information Security Manager
Senior Information Security Manager

Nscale • City Of London

Remote
GBP 110,000 - 140,000
Highly competitive package (base + equity)
Dynamic progression plan tailored to ambitions
Flexible workplace respecting personal lives
+1
Principal Software Engineer - Fleet Management
Principal Software Engineer - Fleet Management

Nscale • Greater London

Remote
GBP 90,000 - 120,000
Competitive salary package
Equity options
Flexible working hours
+2
Senior Manager, Data Centre Vendor Operations & Governance New London
Senior Manager, Data Centre Vendor Operations & Governance New London

Nscale • Greater London

Hybrid
GBP 100,000 - 150,000
Equity options
Health & pension benefits
Career growth
Principal Backbone & Edge Architect
Principal Backbone & Edge Architect

Nscale • Greater London

On-site
GBP 150,000 - 190,000
Equity
Flexible work culture
Diversity & inclusion