Staff Network Engineer

Nscale

Greater London

On-site

GBP 90,000 - 130,000

Full time

15 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Nscale is seeking a Staff Network Engineer to lead and shape the AI-optimized network fabrics, including InfiniBand and Ethernet at scale. You will own architecture, automation, and security across data center sites, guiding complex technical decisions and mentoring engineers.

You will work closely with deployment, data center operations, platform engineering, and vendors to raise the bar on reliability and performance for ML scale-out workloads.

Qualifications

  • 10+ years of network engineering experience in HPC/AI/ hyperscale or large-scale data center environments.
  • Extensive hands-on experience with RDMA-aware networking for AI/HPC workloads, including InfiniBand and RoCE and OpenSM/UFM.
  • Expert-level knowledge of data center routing and control planes: BGP, EVPN-VXLAN, Clos/spine-leaf; production on Cumulus, Nokia, Arista EOS.

Responsibilities

  • Define, design, validate, and evolve large-scale InfiniBand/RoCE and Ethernet fabric architectures at rack, row, and data center scale.
  • Own technical direction for high-performance Ethernet fabrics and establish reference architectures across sites.
  • Design perimeter and security infrastructure (firewalls, NAT, VPN) across WAN and data center edge environments.
  • Lead network automation strategy in a GitOps model with Python/Ansible tooling for provisioning and validation.
  • Drive operational excellence through incident leadership, root-cause analysis, and automation to reduce toil.
  • Set direction for observability, telemetry, monitoring, and alerting of fabric health and performance.

Skills

InfiniBand
RoCE
BGP
EVPN-VXLAN
LACP
QoS
Python
Ansible
Terraform
GitLab CI
GitHub Actions
OpenSM/UFM
Cumulus
Nokia
Arista EOS
Juniper SRX
Palo Alto
Firewall design
Observability
Telemetry

Tools

Python
Ansible
Terraform
GitLab CI
GitHub Actions

Job description

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud strengthens technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

At Nscale, our Engineering team plays a critical role in deploying and operating the infrastructure and software platforms that power our customers.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

About the Role

The Network Engineering Team is responsible for the design, validation, and ongoing operation of all networking services that underpin both the internal management platform and the customer-facing cloud infrastructure — including high-performance Ethernet fabrics, InfiniBand, WAN connectivity, and data center networking. The team also acts as a 3rd/4th line escalation point for the support organization.

As a Staff Network Engineer, you will be a technical authority for Nscale’s AI-optimized network fabrics. You will drive the technical direction across low-latency, high-bandwidth InfiniBand and Ethernet networks supporting large-scale training and inference workloads; own critical technical domains end to end; and raise the bar for architecture, automation, operational rigor, and engineering standards across the organization.

You will combine deep hands‑on engineering with broad architectural influence. You’ll define reference architectures, drive consistency across sites, lead complex technical decisions and escalations, and mentor engineers while partnering closely with deployment, data center operations, platform engineering, and vendors.

What You'll Be Doing
  • Define, design, validate, and evolve large-scale InfiniBand/RoCE and Ethernet fabric architectures at rack, row, and data center scale, with tight integration to bare-metal provisioning and cluster management systems.
  • Own technical direction for high-performance Ethernet fabrics, including BGP, EVPN, VXLAN, LACP, and QoS, and establish reference architectures and standards implemented consistently across sites.
  • Design and engineer perimeter and security infrastructure — firewalls, NAT, VPN, and security policy architecture — across WAN and data center edge environments.
  • Lead network automation strategy in a GitOps model, building and guiding Python/Ansible tooling for provisioning, configuration validation, and compliance, with version-controlled configuration and CI/CD-driven change across multi-vendor environments.
  • Drive operational excellence by leading complex escalations and root-cause analysis for performance and stability issues, and systematically reducing reactive toil through runbooks, automation, and measurable SLOs.
  • Set the direction for network observability, telemetry, monitoring, and alerting to provide clear visibility into fabric health, performance, and traffic patterns.
  • Ensure the accuracy and reliability of source-of-truth network inventory and configuration data, with changes flowing through structured engineering and change-management practices.
  • Partner with deployment, data center operations, platform engineering, systems, storage, and vendors on new site delivery and platform evolution.
  • Act as a technical mentor and force multiplier across the team through architecture reviews, design reviews, incident leadership, documentation, and knowledge sharing.
  • Identify systemic risks and architectural gaps across sites and drive durable solutions that improve scalability, reliability, and operational simplicity.
About You (Skills / Qualifications)
  • 10+ years of network engineering experience, with significant depth in HPC, AI, hyperscale, or large-scale data center environments.
  • Extensive hands‑on experience with RDMA-aware networking for AI/HPC workloads, including InfiniBand and/or RoCE, subnet managers such as OpenSM/UFM, and fabric orchestration.
  • Expert-level knowledge of modern data center routing and control planes, including BGP, EVPN-VXLAN, and Clos/spine-leaf architectures, with production experience on platforms such as Cumulus, Nokia, or Arista EOS.
  • Strong network automation expertise using Python and Ansible, Git-based workflows, and modern infrastructure-as-code and pipeline tooling such as Terraform, GitLab CI, or GitHub Actions; you treat the network as code rather than managing devices by hand.
  • Deep design and engineering experience with firewall platforms such as Juniper SRX and/or Palo Alto, including security policy architecture, high-availability design, and multi‑tenant segmentation.
  • Experience designing network telemetry and observability for high-throughput, performance-sensitive environments.
  • Proven ability to lead complex technical decisions and incidents across networking, systems, storage, and HPC/AI workload teams, with the judgment to balance performance, reliability, operability, and delivery velocity.
  • Demonstrated experience defining architecture, standards, and technical strategy beyond a single project or site, and influencing engineering teams without relying on formal authority.
  • Strong communication and mentoring skills, with the ability to make complex technical trade-offs clear to engineering leaders, operators, and cross‑functional partners.
  • Hands‑on, adaptable, and comfortable operating with high ownership in a fast‑paced environment building next‑generation infrastructure for ML scale‑out.

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enriches our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio‑economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Network Engineer
Principal Network Engineer

Nscale • Greater London

On-site
GBP 120,000 - 170,000
Base + equity
Flexible working
Competitive package
+1
Principal Backbone & Edge Architect
Principal Backbone & Edge Architect

Nscale • Greater London

On-site
GBP 150,000 - 190,000
Equity
Flexible work culture
Diversity & inclusion
(Senior) Infrastructure Engineer (OpenStack Neutron Specialist)
(Senior) Infrastructure Engineer (OpenStack Neutron Specialist)

Nscale • Greater London

On-site
GBP 60,000 - 80,000
Highly competitive compensation package
Performance reviews every 12 months
Dynamic progression plan
+1
Principal Network Platform Lead
Principal Network Platform Lead

Nscale • Greater London

Hybrid
GBP 120,000 - 180,000
Equity
Flexible work environment
Autonomy to shape your day
Principal Network Backbone & Edge Lead
Principal Network Backbone & Edge Lead

Nscale • Greater London

On-site
GBP 120,000 - 180,000
Base + equity
Flexible workplace
Deployment Engineering Director, Systems Engineering
Deployment Engineering Director, Systems Engineering

Nscale • Greater London

On-site
GBP 120,000 - 180,000
Senior Network Engineer
Senior Network Engineer

Radiant • Gloucester

On-site
GBP 100,000 - 130,000
25 days annual leave
Private medical insurance via Bupa
Cycle to Work Scheme
+1
Staff Cloud Native Software Engineer
Staff Cloud Native Software Engineer

Nscale • Greater London

On-site
GBP 110,000 - 170,000
Staff HPC Systems Software Engineer
Staff HPC Systems Software Engineer

AI Startups UK • Greater London

Hybrid
GBP 110,000 - 170,000
Solutions Engineer
Solutions Engineer

Nscale • Greater London

On-site
GBP 90,000 - 120,000