Staff HPC Systems Software Engineer

Nscale

Greater London

On-site

GBP 100,000 - 150,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nscale in London is hiring a Staff HPC Systems Software Engineer to define the technical direction of a core HPC platform. You will scope architecture for Slurm-based services, shaping how multiple teams build, automate, and run the platform in a cloud-native environment.

You’ll work across engineering boundaries to ensure robust, maintainable systems with reusable patterns. The role blends hands-on software engineering with strategic thinking, balancing performance, reliability and operational

Qualifications

  • Extensive experience designing and building production software for HPC systems, especially Slurm-based environments.
  • Strong coding in Go and Python with maintainable, testable design.
  • Ability to define technical direction across multiple teams.
  • Solid understanding of Slurm internals, scheduler behavior and cluster lifecycle.
  • Experience with GPU-backed infrastructure and HPC networking (InfiniBand, RoCE, RDMA).
  • Experience integrating HPC systems with cloud-native platforms and APIs.
  • Experience creating reusable patterns and shared tooling.

Responsibilities

  • Own and evolve the HPC domain's technical direction, including Slurm platform architecture and lifecycle.
  • Define packaging, automation and service exposure for Slurm implementations.
  • Coordinate ownership, interfaces, and operating models across multiple teams.
  • Lead cross-team design for Slurm, Kubernetes-adjacent systems and APIs.
  • Provide hands-on contribution to de-risk critical work.
  • Set patterns for automation, observability and reliability across the HPC platform.
  • Lead initiatives spanning 2–4 teams.
  • Clarify direction to unblock delivery.
  • Influence teams through judgement and design clarity.

Skills

Slurm architecture
Go
Python
Cross-team leadership
System design
GPU HPC networking
Cloud-native integration
Kueue
Communication skills

Tools

Kubernetes

Job description

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performanceinfrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focusedcompanies to achieve superior results by reducing the complexity of AI development. Our GPUcloud bolsters technical capabilities and directly supports strategic business outcomes, includingcost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every teammember takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’llbuild trust through openness and transparency, where everyone is inspired to do their best work. Ifyou join our team, you’ll be contributing to building the technology that powers the future.

About The Role

We’re hiring a Staff HPC Systems Software Engineer to definethe technical direction and evolution of a core HPC platform domain at Nscale. In this role, you will operate beyond a single team, shaping how multiple teams build, automate,and run Slurm-based capabilities within Nscale’s wider cloud-native platform. You’ll work acrossengineering boundaries to bring coherence to architecture, interfaces, lifecycle models, andoperational approaches, while partnering closely with teams working on platform tooling,infrastructure APIs, identity systems, and Kubernetes-adjacent systems.

This is a high-impact staff-level role for someone who combines deep hands-on softwareengineering with strong systems judgement. Your work will help ensure Nscale’s HPC services arerobust, supportable, and maintainable, while creating leverage through shared patterns, reusableimplementations, and clear technical direction across ambiguous, business-critical problemspaces.

What You’ll Be Doing
Domain Architecture & Technical Direction
  • Own and evolve the technical direction for a defined HPC systems domain, such as Slurmplatform architecture, scheduler integrations, cluster lifecycle, workload environments orservice automation.
  • Make architectural decisions that balance software quality, operational realities, customerneeds, and long-term maintainability.
  • Define how proven Slurm implementations should be packaged, automated and exposedas a service.
  • Resolve ambiguity around ownership, interfaces, lifecycle boundaries, and operatingmodels across teams.
  • Act as the technical escalation point for the most complex issues within the domain.
Cross-Team Engineering Leverage
  • Establish shared patterns for automation, service lifecycle management, observability,reliability and supportability across the HPC platform.
  • Drive cross-team design for integrations between Slurm, Kubernetes-adjacent systems,infrastructure APIs, identity systems and platform tooling.
  • Create reusable modules, automation, deployment patterns, and referenceimplementations that increase engineering leverage.
  • Identity and correct avoidable technical divergence, duplicated effort and fragile operatingmodels
  • Ensure domain designs reflect the realities of GPU scheduling, HPC networking, performance isolation and production operations.
Delivery, Reliability & Influence
  • Lead technically critical initiatives spanning 2-4 teams or a defined HPC platform area.
  • Unblocked delivery by clarifying technical direction and reducing ambiguity in complexsystem design problems.
  • Contribute hands-on where needed to de-risk or accelerate critical work.
  • Influence engineering teams without formal authority through strong judgement, designclarity and practical solutions.
  • Partner with adjacent cloud-native software engineers so HPC implementations build onshared platform patterns rather than separate ones.
KPIs
  • Technical direction across a defined HPC domain
  • Delivery of critical initiatives across 2-4 teams
  • Reduction in technical divergence and duplicated effort
  • Reliability and supportability of Slurm-based HPC services
About You
  • Extensive experience designing and building production software and automation for HPCsystems, especially Slurm-based environments.
  • Strong track record of writing maintainable, testable and resilient software in Go, Python orsimilar languages.
  • Proven ability to define technical direction across a domain spanning multiple teams orservices.
  • Strong understanding of Slurm internals, scheduler behaviour, cluster lifecycle concernsand operational trade offs.
  • Strong practical understanding of GPU-backed infrastructure and HPC networking,including InfiniBand, RoCE, RDMA and performance sensitive workload characteristics.
  • Experience integrating HPC systems with cloud-native platforms, APIs, or service deliverymodels.
  • Experience creating engineering leverage through standards, reusable patterns, sharedtooling and architectural clarity.
  • Strong judgement in balancing short-term delivery with long-term platform health andsupportability.
  • Strong written and verbal communication skills, with the ability to align multiple teamsaround a coherent technical direction.
  • Experience with other schedulers or batch systems such as Kueue is valuable.
What We Can Offer You

You’ll have the opportunity to help shape the operating standards behind a next-generation AI cloud platform, working on complex infrastructure challenges with real ownership and impact. This is a chance to play a meaningful role in scaling high-performance, sustainable data centre operations in a fast-moving environment

Equal Opportunities Statement

We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there’s anything we can do to accommodate your specific situation, please let us know. The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Staff Security Engineer, Platform Security
Sr. Staff Security Engineer, Platform Security

Uncover • Greater London

Hybrid
GBP 155,000 - 185,000
Principal Software Engineer - Fleet Management
Principal Software Engineer - Fleet Management

Nscale • Greater London

Remote
GBP 90,000 - 120,000
Competitive salary package
Equity options
Flexible working hours
+2
(Senior) Infrastructure Engineer (OpenStack Neutron Specialist)
(Senior) Infrastructure Engineer (OpenStack Neutron Specialist)

Nscale • Greater London

On-site
GBP 60,000 - 80,000
Highly competitive compensation package
Performance reviews every 12 months
Dynamic progression plan
+1
(Senior) Infrastructure Engineer (OpenStack Ironic Specialist)
(Senior) Infrastructure Engineer (OpenStack Ironic Specialist)

Nscale • Greater London

On-site
GBP 50,000 - 70,000
Highly competitive compensation package
Performance reviews every 12 months
Flexible paid time off
+2
Principal Network Engineer
Principal Network Engineer

AI Startups UK • Greater London

Hybrid
GBP 120,000 - 190,000
Base + equity
Principal Network Platform Lead
Principal Network Platform Lead

AI Startups UK • Greater London

Hybrid
GBP 110,000 - 140,000
competitive compensation
Base + equity
Flexible work options
Principal Engineer, Storage Services
Principal Engineer, Storage Services

Nscale • Greater London

On-site
GBP 90,000 - 130,000
Competitive compensation package including base, bonus, and equity
Flexible working environment
Performance reviews every 12 months
Principal Network Engineer
Principal Network Engineer

Nscale • Greater London

On-site
GBP 120,000 - 170,000
Base + equity
Flexible working
Competitive package
+1
Solutions Engineer
Solutions Engineer

Nscale • Greater London

On-site
GBP 90,000 - 120,000
Principal Network Platform Lead
Principal Network Platform Lead

Nscale • Greater London

Hybrid
GBP 120,000 - 180,000
Equity
Flexible work environment
Autonomy to shape your day