Senior Infrastructure Operations Engineer (EMEA)

United States Digital Space LLC

Greater London

Hybrid

GBP 116,000 - 128,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health coverage
Equity RSUs
401(k) matching
Unlimited PTO
Hybrid work model
Office meals

Job summary

Lightning AI in London is seeking an experienced Senior Infrastructure Operations Engineer to help operate and scale the infrastructure behind one of the world's largest GPU fleets. The role sits at the center of production reliability across Linux, bare-metal servers, GPUs, networking, storage, provisioning, and cluster orchestration.

You will own provisioning, monitoring, diagnostics, and on-call support, build automation, and partner with Infrastructure Engineering and Platform teams to

Qualifications

  • Strong Linux production infrastructure experience at scale.
  • Experience troubleshooting complex infra across compute, networking, storage, and OS.
  • Strong automation and scripting with Python, Go, Bash, Ansible, or similar.
  • Experience with Kubernetes, Slurm, or other cluster/workload schedulers.
  • Experience with monitoring, observability, and telemetry to diagnose issues.
  • Strong systems and networking fundamentals with ownership of production issues.
  • Comfortable collaborating across engineering, operations, and customer-facing teams.

Responsibilities

  • Operate and troubleshoot large-scale GPU and bare-metal infrastructure across Linux, compute, networking, storage, and cluster orchestration.
  • Serve as a technical escalation point for complex infrastructure issues, driving resolution and root cause analysis.
  • Own provisioning, configuration, validation, maintenance, remediation, and decommissioning workflows.
  • Build automation and internal tooling to scale the infrastructure fleet.
  • Identify recurring failures and work with Infra Eng to improve reliability and automation.
  • Improve provisioning, monitoring, diagnostics, and operational processes for a growing footprint.
  • Collaborate with Infrastructure Engineering, Network Engineering, Data Center Operations, Customer Experience, and Platform Engineering to resolve cross-team issues.
  • Participate in a distributed primary/secondary on-call rotation.

Skills

Linux ops
Automation scripts
Python scripting
Go scripting
Ansible
Kubernetes
Slurm
Monitoring
Telemetry
Networking basics
On-call readiness

Tools

PXE
BMC
IPMI
Redfish
iDRAC
NVIDIA GPUs
DCGM
InfiniBand
RoCE/RDMA
GPFS
Ceph

Job description

Who We Are

Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.

Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.

We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.

The Way We Work

The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:

What We're Looking For

Lightning AI is seeking an experienced Senior Infrastructure Operations Engineer to help operate and scale the infrastructure behind one of the world's largest GPU fleets.

Our Infrastructure Operations team sits at the center of production reliability for Lightning's AI infrastructure. The team works across Linux, bare-metal servers, GPUs, networking, storage, provisioning, and cluster orchestration, serving as a critical technical escalation point when complex infrastructure issues arise.

What You'll Do
  • Operate and troubleshoot large-scale GPU and bare-metal infrastructure across Linux, compute, networking, storage, and cluster orchestration.
  • Serve as a technical escalation point for complex infrastructure issues, driving problems from initial investigation through resolution and root cause analysis.
  • Own operational workflows across the infrastructure lifecycle, including provisioning, configuration, validation, maintenance, remediation, and decommissioning.
  • Build automation and internal tooling that eliminates repetitive operational work and enables the infrastructure fleet to scale efficiently.
  • Identify recurring failure modes and partner with Infrastructure Engineering to build more reliable, repeatable, and automated systems.
  • Improve provisioning, monitoring, diagnostics, and operational processes across a rapidly growing infrastructure footprint.
  • Partner closely with Infrastructure Engineering, Network Engineering, Data Center Operations, Customer Experience, and Platform Engineering to resolve issues that cross team or system boundaries.
  • Participate in a distributed primary/secondary on-call rotation supporting production infrastructure.
What You'll Need
Required Qualifications
  • Strong experience operating and troubleshooting Linux-based production infrastructure at scale.
  • Experience troubleshooting complex infrastructure issues across compute, networking, storage, and operating systems.
  • Strong automation and scripting skills using Python, Go, Bash, Ansible, or similar tools.
  • Experience with Kubernetes, Slurm, or other cluster and workload orchestration systems.
  • Experience using monitoring, observability, and telemetry systems to diagnose and troubleshoot production infrastructure.
  • Strong systems and networking fundamentals, with a track record of owning production issues through resolution.
  • Comfortable working in ambiguous environments and collaborating across engineering, operations, and customer-facing teams.
Nice-to-Haves
  • Experience troubleshooting, provisioning, or operating bare-metal server infrastructure.
  • Experience operating or troubleshooting GPU, HPC, or other high-performance compute infrastructure.
  • Familiarity with NVIDIA GPUs, DCGM, InfiniBand, RoCE/RDMA, NVLink, or high-speed data center networking.
  • Experience with hardware management and provisioning technologies such as PXE, BMC, IPMI, Redfish, or iDRAC.
  • Experience with distributed or high-performance storage systems such as VAST, Ceph, GPFS, or WEKA.
  • Experience with infrastructure-as-code, configuration management, or GitOps workflows.
Compensation

We are committed to offering competitive compensation that reflects the value each team member brings to our mission. Final offers are based on factors such as experience, skills, geographic location, and role expectations. In addition to base salary, our total rewards package for eligible roles includes a discretionary bonus, a meaningful equity component, and comprehensive benefits.

The anticipated annual base salary range for this role is:

£116,000—£128,000 GBP

Benefits and Perks

We offer a comprehensive and competitive benefits package designed to support our employees’ health, well-being, and long-term success:

  • Comprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.
  • Meaningful Equity: RSUs that give employees a stake in the company's long-term success.
  • Retirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).
  • Flexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.
  • Company-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.
  • Paid Parental & Family Leave: Paid leave to support you and your family through life's important moments.
  • Professional Development: Annual learning and development allowance to support your professional growth.
  • Wellness Benefits: Wellness and work-from-home stipends to support your physical and mental well-being.
  • Sabbatical Program: Four weeks of paid sabbatical leave after four years of service.
  • Flexible Work: Flexible schedules and a hybrid work model for our office-based teams.
  • In-Office Meals: Complimentary meals at our office hubs.

Benefits may vary by location, team, and role.

*At Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.*

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Infrastructure Operations Engineer (EMEA) New London, England, United Kingdom
Senior Infrastructure Operations Engineer (EMEA) New London, England, United Kingdom

Lightning • Greater London

Hybrid
GBP 90,000 - 150,000
Health coverage
Equity
401(k) matching
+3
Senior Infrastructure Software Engineer
Senior Infrastructure Software Engineer

Lightningai • Greater London

On-site
GBP 133,000 - 163,000
Health coverage
Equity
401k matching
+8
Senior Infrastructure Software Engineer
Senior Infrastructure Software Engineer

Precision Labs • Greater London

Hybrid
GBP 134,000 - 164,000
Comprehensive Health Coverage
Meaningful Equity
401(k) Matching (US)
+7
Senior Infrastructure Software Engineer New York, New York, United States
Senior Infrastructure Software Engineer New York, New York, United States

Lightning • Greater London

Hybrid
GBP 134,000 - 164,000
Health coverage
Equity: meaningful RSUs
401(k) retirement matching
+4
AI Platform Support Engineer (EMEA)
AI Platform Support Engineer (EMEA)

Lightningai • Greater London

On-site
GBP 75,000 - 95,000
Comprehensive health coverage
Equity (RSUs)
Hybrid work model
+2
Senior Software Engineer, Core Platform
Senior Software Engineer, Core Platform

Lightningai • Greater London

On-site
USD 180,000 - 250,000
Comprehensive Health Coverage
Equity (RSUs)
401(k) matching
+9
Senior Infrastructure Operations Engineer (EMEA)
Senior Infrastructure Operations Engineer (EMEA)

lightningai • Greater London

Hybrid
GBP 90,000 - 130,000
Senior Infrastructure Operations Engineer (EMEA)
Senior Infrastructure Operations Engineer (EMEA)

Academic Key • Greater London

Hybrid
GBP 110,000 - 155,000
Senior GPU Infra Operations Engineer – EMEA
Senior GPU Infra Operations Engineer – EMEA

Lightning • Greater London

Hybrid
GBP 90,000 - 150,000
Health coverage
Equity
401(k) matching
+3
Infrastructure Site Reliability Engineer
Infrastructure Site Reliability Engineer

Radiant • Gloucester

On-site
GBP 70,000 - 110,000
25 days annual leave
Cycle to Work Scheme
Gympass subscription