Remote Senior HPC Cluster Architect for DL Compute & Storage

NVIDIA Corporation

Warszawa

On-site

PLN 221,250 - 507,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA Corporation in Poland (Warsaw) is seeking a Senior HPC Cluster Administrator to lead design, deployment, and reliability of our large GPU compute clusters across DGX, HGX, and Grace platforms. You will own the full lifecycle, design storage strategies, automate with Ansible and Terraform, manage Slurm scheduling, and monitor systems with Prometheus and Grafana.

Collaboration with ML teams to optimize distributed training is essential.

Qualifications

  • BS/MS in CS, EE, CE, or equivalent hands-on experience.
  • 5+ years deploying and administering large-scale HPC or ML training clusters.
  • Strong Linux systems administration at scale.
  • Scripting in Python and/or Bash.
  • Experience with Slurm scheduling and accounting.
  • Configuration management and IaC tools (Ansible required).
  • Experience with containers (Docker, Apptainer/Singularity, Kubernetes).
  • High-speed networking (InfiniBand, RoCE, RDMA, EFA).
  • Distributed/parallel filesystems and storage architectures.

Responsibilities

  • Own the full lifecycle of GPU compute clusters — procurement, provisioning, configuration, monitoring, and deprecation.
  • Design and scale storage solutions with capacity and performance roadmaps.
  • Lead infrastructure automation using Ansible, Terraform and CI/CD pipelines.
  • Manage and optimize Slurm scheduling, fair-share policies, and MIG/GPU partitioning.
  • Maintain observability stacks (Prometheus, Grafana, DCGM) and resolve incidents.
  • Collaborate with ML engineers to tune cluster configs for large-scale training.
  • Evaluate new tech to improve performance and reliability.

Skills

Linux administration
Python
bash scripting
Slurm
Ansible
Terraform
Docker
Kubernetes
InfiniBand
RoCE RDMA EFA

Education

BS/MS in CS/EE/CE

Tools

Ansible
Terraform

Job description

NVIDIA Corporation in Poland (Warsaw) is seeking a Senior HPC Cluster Administrator to lead design, deployment, and reliability of our large GPU compute clusters across DGX, HGX, and Grace platforms. You will own the full lifecycle, design storage strategies, automate with Ansible and Terraform, manage Slurm scheduling, and monitor systems with Prometheus and Grafana.

Collaboration with ML teams to optimize distributed training is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Cluster Architect for Deep Learning Infra
Senior HPC Cluster Architect for Deep Learning Infra

NVIDIA Gruppe • Warszawa

On-site
PLN 221,000 - 507,000
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA Gruppe • Warszawa

On-site
PLN 221,000 - 507,000
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA Corporation • Warszawa

On-site
PLN 221,000 - 507,000
Senior AI Network Architect for Scalable GPU Infrastructure
Senior AI Network Architect for Scalable GPU Infrastructure

NVIDIA • Poland

On-site
PLN 221,000 - 507,000
Competitive salary
Comprehensive benefits package
Benefits portal access
Senior Forward Deployed Solution Engineer (Poland) Engineering · Wroclaw · Onsite
Senior Forward Deployed Solution Engineer (Poland) Engineering · Wroclaw · Onsite

Spectro Cloud • Wrocław

On-site
PLN 180,000 - 300,000
Senior Forward Deployed Solution Engineer (Poland)
Senior Forward Deployed Solution Engineer (Poland)

Spectro Cloud • Wrocław

On-site
PLN 260,000 - 380,000
Senior Software Engineer, Cloud Automation
Senior Software Engineer, Cloud Automation

NVIDIA Gruppe • Warszawa

On-site
PLN 183,000 - 416,000
Senior HPC and AI Network Software Architect
Senior HPC and AI Network Software Architect

NVIDIA • Poland

On-site
PLN 221,000 - 507,000
Competitive salary
Comprehensive benefits package
Benefits portal access
Senior Software Engineer, Cloud Automation
Senior Software Engineer, Cloud Automation

NVIDIA • Warszawa

On-site
PLN 183,000 - 416,000
Senior Cloud Automation Engineer - AI-Driven Ops
Senior Cloud Automation Engineer - AI-Driven Ops

NVIDIA Gruppe • Warszawa

On-site
PLN 183,000 - 416,000