HPC Platform & Slurm Lead Engineer - Hybrid, Sydney

NTT DATA, Inc.

Sydney

Hybrid

AUD 150,000 - 190,000

Full time

38 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

NTT DATA, Inc. seeks an HPC Platform & Slurm Lead Engineer to design, deploy and operate a large-scale HPC environment in Sydney. The role focuses on Slurm administration, GPU infrastructure, Linux engineering and automation, with hybrid work arrangements available.

The successful candidate will lead workload onboarding, performance optimisation and platform reliability across AI/ML workloads and scientific computing. ASAP start and potential conversion to permanent.

Qualifications

  • Proven experience designing and operating HPC platforms in production.
  • Deep expertise with Slurm Workload Manager and related tooling (SlurmDBD, accounting, partitions).
  • Experience with NVIDIA GPU infrastructure and AI/ML workloads.
  • Familiarity with research/university or large enterprise HPC environments is valued.

Responsibilities

  • Design, deploy and support enterprise-scale HPC environments using Slurm Workload Manager.
  • Architect and administer highly available Slurm controller infrastructure, scheduling policies, QoS and fairshare models.
  • Build, configure and maintain Linux-based compute, login and management nodes.
  • Deploy and support NVIDIA GPU platforms, including CUDA, NCCL, DCGM and GPU scheduling.

Skills

HPC platform design
Slurm administration
Linux engineering
GPU infrastructure
Automation scripting
Cluster management

Tools

Grafana
Prometheus
DCGM
CUDA
NCCL

Job description

NTT DATA, Inc. seeks an HPC Platform & Slurm Lead Engineer to design, deploy and operate a large-scale HPC environment in Sydney. The role focuses on Slurm administration, GPU infrastructure, Linux engineering and automation, with hybrid work arrangements available.

The successful candidate will lead workload onboarding, performance optimisation and platform reliability across AI/ML workloads and scientific computing. ASAP start and potential conversion to permanent.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Platform & Slurm Lead Engineer
HPC Platform & Slurm Lead Engineer

NTT DATA, Inc. • Sydney

Hybrid
AUD 150,000 - 190,000
Senior HPC Systems Engineer - Linux, Slurm & Networking
Senior HPC Systems Engineer - Linux, Slurm & Networking

Jump Trading • Sydney

On-site
AUD 140,000 - 190,000
Platform & Distributed Systems Engineer — Scalable HPC
Platform & Distributed Systems Engineer — Scalable HPC

Tribus • Sydney

On-site
AUD 100,000 - 150,000
Linux HPC Operations Engineer - 24/7 High-Performance Infra
Linux HPC Operations Engineer - 24/7 High-Performance Infra

Westbury Partners • Sydney

On-site
AUD 120,000 - 170,000
Solutions Architect - Systems Integrator
Solutions Architect - Systems Integrator

Hamilton Barnes • Sydney

On-site
AUD 140,000 - 210,000
Field HPC Infrastructure Engineer
Field HPC Infrastructure Engineer

Nityo Infotech • City of Melbourne

On-site
AUD 90,000 - 120,000
HPC Systems Engineer
HPC Systems Engineer

Jump Trading • Sydney

On-site
AUD 140,000 - 190,000
Senior Infrastructure Engineer - HPC & Storage (Hybrid)
Senior Infrastructure Engineer - HPC & Storage (Hybrid)

Experis Executive • City of Melbourne

Hybrid
AUD 120,000 - 180,000
Weekly pay
Hybrid working arrangements
Latest cutting-edge technology
Site Reliability Engineer: AI Infra & HPC
Site Reliability Engineer: AI Infra & HPC

Firmus Technologies • City of Melbourne

On-site
AUD 120,000 - 180,000
HPC Systems Engineer
HPC Systems Engineer

P2P • Sydney

On-site
AUD 140,000 - 190,000