Staff HPC Systems Software Engineer

Nscale

New York (NY)

On-site

USD 180,000 - 260,000

Full time

13 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Bonus
Equity
Medical Insurance
Dental Insurance
Vision Insurance
Flexible Paid Time Off
Parental Leave
Retirement Plan Participation

Job summary

Nscale is seeking a Senior HPC Platform Architect to define the technical direction and architecture for its Slurm-based HPC platform. You will lead cross-team initiatives to integrate HPC systems with cloud-native APIs, driving reliability, scalability, and performance.

The role requires extensive production experience with Slurm in HPC environments using Go or Python, deep GPU infrastructure knowledge, and strong leadership across multiple engineering teams.

Qualifications

  • Extensive production software experience in Slurm-based HPC environments using Go or Python.
  • Deep knowledge of GPU-backed infrastructure and HPC networking (InfiniBand/RDMA).
  • Ability to lead technical strategy across multiple teams.

Responsibilities

  • Define the technical direction and architecture for Nscale's Slurm-based HPC platform domain.
  • Lead cross-team initiatives to integrate HPC systems with cloud-native APIs.
  • Improve system reliability and scalability across the HPC stack.

Skills

Slurm
Go
Python
HPC Systems Architecture
GPU Scheduling
InfiniBand
RoCE
RDMA
Kubernetes
Cloud-Native Platforms
Service Automation
Systems Software Engineering

Job description

Define the technical direction and architecture for Nscale's Slurm-based HPC platform domain. Lead cross-team initiatives to integrate HPC systems with cloud-native APIs and improve overall system reliability and scalability.

Requirements:

Extensive experience building production software for Slurm-based HPC environments using Go or Python. Deep knowledge of GPU-backed infrastructure, HPC networking (InfiniBand/RDMA), and the ability to lead technical strategy across multiple teams.

Key Skills:
  • Slurm
  • Go
  • Python
  • HPC Systems Architecture
  • GPU Scheduling
  • InfiniBand
  • RoCE
  • RDMA
  • Kubernetes
  • Cloud-Native Platforms
  • Service Automation
  • Systems Software Engineering
Benefits:
  • Bonus
  • Equity
  • Medical Insurance
  • Dental Insurance
  • Vision Insurance
  • Flexible Paid Time Off
  • Parental Leave
  • Retirement Plan Participation
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Systems Engineer — Slurm, GPU, Cloud-Native
Senior HPC Systems Engineer — Slurm, GPU, Cloud-Native

Nscale • New York (NY)

On-site
USD 180,000 - 260,000
Bonus
Equity
Medical Insurance
+5
Software Engineer, HPC Scheduling
Software Engineer, HPC Scheduling

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 90,000 - 120,000
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Senior Software Engineer, Cloud-Native Stack – CSP Engagements

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Director, HPC Systems Software Engineering
Director, HPC Systems Software Engineering

Nscale • New York (NY)

On-site
USD 230,000 - 343,000
Medical insurance
Dental & Vision
Flexible paid time off
+1
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Austin (TX)

On-site
USD 152,000 - 242,000
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Westford (MA)

On-site
USD 152,000 - 242,000
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Durham (NC)

On-site
USD 152,000 - 242,000
Director, HPC Systems Software Engineering
Director, HPC Systems Software Engineering

Nscale • Seattle (WA)

On-site
USD 230,000 - 343,000
Medical benefits
Retirement plan
Parental leave
+1
Senior Systems Engineer
Senior Systems Engineer

CoreWeave • Bellevue (WA)

On-site
USD 140,000 - 190,000
Medical Insurance
401(k) with Employer Match
Flexible PTO
+1
Staff Slurm Cluster & HPC Engineer
Staff Slurm Cluster & HPC Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 250,000