HPC Platform & Slurm Lead Engineer

NTT DATA, Inc.

Sydney

Hybrid

AUD 150,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NTT DATA, Inc. seeks an HPC Platform & Slurm Lead Engineer to design, deploy and operate a large-scale HPC environment in Sydney. The role focuses on Slurm administration, GPU infrastructure, Linux engineering and automation, with hybrid work arrangements available.

The successful candidate will lead workload onboarding, performance optimisation and platform reliability across AI/ML workloads and scientific computing. ASAP start and potential conversion to permanent.

Qualifications

  • Proven experience designing and operating HPC platforms in production.
  • Deep expertise with Slurm Workload Manager and related tooling (SlurmDBD, accounting, partitions).
  • Experience with NVIDIA GPU infrastructure and AI/ML workloads.
  • Familiarity with research/university or large enterprise HPC environments is valued.

Responsibilities

  • Design, deploy and support enterprise-scale HPC environments using Slurm Workload Manager.
  • Architect and administer highly available Slurm controller infrastructure, scheduling policies, QoS and fairshare models.
  • Build, configure and maintain Linux-based compute, login and management nodes.
  • Deploy and support NVIDIA GPU platforms, including CUDA, NCCL, DCGM and GPU scheduling.

Skills

HPC platform design
Slurm administration
Linux engineering
GPU infrastructure
Automation scripting
Cluster management

Tools

Grafana
Prometheus
DCGM
CUDA
NCCL

Job description

Role Title: HPC Platform & Slurm Lead Engineer

Location: Sydney, NSW (Hybrid)

Start Date: ASAP

Duration: 12-Month Contract (with potential permanent conversion)

Working Flexibility: Hybrid

NTT DATA is seeking a highly skilled HPC Platform & Slurm Lead Engineer to lead the design, deployment, optimisation and operational management of a large-scale High Performance Computing (HPC) environment. This role will play a critical part in delivering a production-grade research computing platform supporting AI, machine learning, data analytics and advanced scientific workloads. The successful candidate will provide hands-on technical leadership across Slurm administration, Linux engineering, GPU infrastructure, high-performance storage, automation and platform operations.

Is innovation part of your DNA? Do you want to enable a connected future for people, organizations, and society?

Join our growing global NTT team and you'll be part of the world's largest ICT company (by revenue). We've combined the capabilities of 28 remarkable companies to become one, leading technology services provider. Together, we help our people, clients, and communities do great things with technology to create a more secure and connected future. We employ 40,000 people across 57 countries. By bringing together the world's best technology companies and emerging innovators, we work together to deliver sustainable outcomes to businesses and the world. Innovation is part of our DNA. We believe it's key to what makes us different. So, we strive to move forward, challenge the status quo, and drive excellence through the technologies we integrate and the services we deliver around the world. The result is connected cities, connected factories, connected healthcare, connected agriculture, connected conservation, connected mobility, and connected sport. Together we enable the connected future.

Key Responsibilities
  • Design, deploy and support enterprise-scale HPC environments using Slurm Workload Manager.
  • Architect and administer highly available Slurm controller infrastructure, scheduling policies, QoS frameworks and fairshare models.
  • Build, configure and maintain Linux-based compute, login and management nodes.
  • Deploy and support NVIDIA GPU platforms, including CUDA, NCCL, DCGM and GPU scheduling capabilities.
  • Assist research and engineering teams with workload onboarding, performance optimisation and resource utilisation.
  • Implement automation, monitoring and observability solutions using tools such as Grafana and Prometheus.
  • Drive platform reliability, security, patching, operational processes, reporting and continuous improvement initiatives.
Skills & Experience
  • Strong experience designing, implementing and supporting HPC platforms in production environments.
  • Deep expertise administering Slurm Workload Manager, SlurmDBD, accounting, reservations, partitions, QoS and scheduler optimisation.
  • Experience supporting NVIDIA GPU infrastructure, AI/ML workloads and large-scale compute environments.
  • Knowledge of HPC storage, container technologies, high-performance networking and cluster management solutions.
  • Strong scripting, automation and infrastructure lifecycle management capabilities.
  • Experience within research, university, scientific computing or large enterprise environments will be highly regarded.

Questions? Reach out: kumar.kuchipudi@global.ntt

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Platform & Slurm Lead Engineer - Hybrid, Sydney
HPC Platform & Slurm Lead Engineer - Hybrid, Sydney

NTT DATA, Inc. • Sydney

Hybrid
AUD 150,000 - 190,000
Senior HPC Systems Engineer - Linux, Slurm & Networking
Senior HPC Systems Engineer - Linux, Slurm & Networking

Jump Trading • Sydney

On-site
AUD 140,000 - 190,000
HPC Systems Engineer
HPC Systems Engineer

P2P • Sydney

On-site
AUD 140,000 - 190,000
Site Reliability Engineer, AI Infrastructure
Site Reliability Engineer, AI Infrastructure

Firmus Technologies • City of Melbourne

On-site
AUD 120,000 - 180,000
Senior HPC Systems Engineer - Linux & Scale
Senior HPC Systems Engineer - Linux & Scale

P2P • Sydney

On-site
AUD 140,000 - 190,000
Senior Networking Solutions Architect - AI/HPC Cloud
Senior Networking Solutions Architect - AI/HPC Cloud

NVIDIA • Sydney

On-site
AUD 120,000 - 160,000
Solutions Architect - Systems Integrator
Solutions Architect - Systems Integrator

Hamilton Barnes • Sydney

On-site
AUD 140,000 - 210,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Experis • City of Melbourne

Hybrid
AUD 140,000 - 200,000
Weekly pay
Hybrid work arrangements
Cutting-edge technology access
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Experis Australia • City of Melbourne

Hybrid
AUD 120,000 - 180,000
Weekly pay
Hybrid working arrangements
Work with cutting edge technology
+2
HPC Systems Engineer
HPC Systems Engineer

Jump Trading • Sydney

On-site
AUD 140,000 - 190,000