Senior Manager, Validation and HPC - NVIS

NVIDIA

Holmdel Township (NJ)

On-site

USD 216,000 - 397,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity and benefits

Job summary

NVIDIA seeks an HPC Deployment Manager to lead the deployment of cutting-edge InfiniBand and Ethernet technologies for AI HPC systems. You will supervise a team of AI and HPC experts, guiding planning, implementation, and validation across hardware and software platforms.

Responsibilities include coordinating multi-vendor hardware and software deployment, writing procedures for validation, and ensuring operational reliability.

Qualifications

  • 10+ years in IT/HPC or related field.
  • 3+ years in a management/leadership role.
  • Deep knowledge of HPC systems design, configuration, and planning.
  • Solid HPC storage knowledge.
  • Proficiency with low latency/high-bandwidth interconnects (InfiniBand, Ethernet).
  • Experience with HPC cluster management/provisioning tools and job schedulers (Slurm, Salt, xCAT).
  • Proficiency with distributed memory parallelism (OpenMP, MPI, NCCL, HPL) and GPUs.
  • Strong scripting (Bash, Perl, Python).
  • Administration of secure Linux/Unix systems (CentOS, Solaris).
  • Ability to manage large, complex systems and troubleshoot network issues.

Responsibilities

  • Directs and supervises HPC engineering funtions in design, install, and validation of AI HPC systems.
  • Leads planning, implementation, and performance of HPC projects; enhances system service brings-up and maintenance.
  • Drives hardware/software deployment plans and procedures for system validation.
  • Leads team activities, tests, and plans for AI HPC implementations, scripts, and testing procedures.
  • Collaborates with internal teams to deliver quality service and drive continuous improvement.
  • Mentors team members and supports career growth.

Skills

HPC leadership
Interconnect infra
Slurm
OpenMP
MPI
Python
Bash
Linux admin
Multi-vendor hardware

Education

Bachelor's degree in CS or related field
Equivalent experience

Tools

Ansible
Salt
Puppet

Job description

NVIDIA is in search of an HPC Deployment Manager to bolster our NVIDIA Infrastructure Specialists division! Across academia and industry, NVIDIA's products are driving ground-breaking advancements in deep learning, data analytics, and the optimization of data centers. Join our team, where we are at the forefront of constructing some of the globe's most expansive and rapid data centers! We seek an individual capable of supervising the deployment of cutting-edge InfiniBand and Ethernet technologies with a team comprising AI and HPC experts. This role demands dynamic interpersonal abilities and a customer-centric approach.

What You Will Be Doing
  • Directs and supervises the service HPC engineering functions in designing, developing, installing, and validating hardware and software for the Customer AI High-Performance Computing (HPC) systems.
  • Responsible for leading our HPC projects' planning, implementation, and performance. Improves the integrity of system services bring-up and related by applying groundbreaking technical and operational knowledge to configure and maintain HPC AI network and server platforms.
  • Drives HPC team hardware and software deployment, plans, develops, and deploys procedures for system validation.
  • Lead team activities and drive tests and plans for Customer's HPC AI systems implementations, custom scripts, and testing procedures to ensure operational reliability for the system.
  • Supports the HPC Engineering team, working with other internal collaborators to develop and run a well-rounded strategy for delivering service quality and continuous service improvement.
  • Leads team member development, helping them set and achieve goals for their career growth.
  • Build strong relationships with NVIDIA leaders, customers, partners, and collaborators. Works closely to identify, implement, and support leading NVIDIA's AI solutions engineering, maintaining currency with industry standards and innovations.
  • Be the domain authority with customers during planning calls through implementation.
What We Need To See
  • 10+ overall years' experience in IT, high-performance computing, or other related field; 3+ years of experience in a management or leadership role
  • Demonstrated expertise in HPC systems design configuration and planning, and solid knowledge of HPC storage
  • Proficiency with low latency/high-bandwidth interconnect infrastructure (Infiniband and Ethernet).
  • Expertise with HPC system software cluster management/provisioning tools, including job schedulers (Slurm, salt, xCAT).
  • Proficiency with shared and distributed memory parallelism (OpenMP, MPI, NCCL and HPL) and accelerators (GPUs).
  • Strong scripting ability (Bash, Perl, Python, etc.) and experience with programming fundamentals.
  • Expertise with administration, supervising and maintaining secure Linux/Unix operating systems (CentOS, Solaris).
  • Ability to understand and work with large, sophisticated systems, identify and resolve problems, handle performance, and troubleshoot network issues related to infrastructure.
  • Expertise with multi-vendor hardware/software management, security, and network/Internet protocols
  • Bachelor's degree in computer science, information systems, or a related field or equivalent experience
Ways To Stand Out From The Crowd
  • InfiniBand experience.
  • Experience with GPU-focused hardware/software.
  • Experience with MPI.
  • Automation tooling background (Ansible, Salt, Puppet, etc.).
  • Ethernet and Storage technologies such as Lustre or GPFS.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 216,000 USD - 345,000 USD for Level 4, and 248,000 USD - 396,750 USD for Level 5.

You will also be eligible for equity and benefits.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, Validation and HPC - NVIS
Senior Manager, Validation and HPC - NVIS

NVIDIA • United States

On-site
USD 216,000 - 397,000
Equity
Senior Manager, Validation and HPC - NVIS
Senior Manager, Validation and HPC - NVIS

NVIDIA • Washington

On-site
USD 216,000 - 397,000
Equity
Benefits
Senior Manager, Validation and HPC - NVIS
Senior Manager, Validation and HPC - NVIS

NVIDIA • Austin (TX)

On-site
USD 216,000 - 397,000
Equity
Benefits
Senior Manager, Validation and HPC - NVIS
Senior Manager, Validation and HPC - NVIS

NVIDIA • South Carolina

On-site
USD 230,000 - 360,000
Senior Manager, Validation and HPC - NVIS
Senior Manager, Validation and HPC - NVIS

Nvidia Corporation • Austin (TX)

On-site
USD 216,000 - 397,000
Equity
Benefits
Senior Manager, Validation and HPC - NVIS
Senior Manager, Validation and HPC - NVIS

NVIDIA • Virginia (MN)

On-site
USD 216,000 - 397,000
Equity
Benefits
Senior Manager, Validation and HPC - NVIS
Senior Manager, Validation and HPC - NVIS

NVIDIA • North Carolina

On-site
USD 216,000 - 397,000
Equity
Benefits
Senior Manager, Validation and HPC - NVIS
Senior Manager, Validation and HPC - NVIS

NVIDIA AI • Holmdel Township (NJ)

On-site
USD 216,000 - 397,000
Equity
Benefits
Senior Manager, Validation and HPC - NVIS
Senior Manager, Validation and HPC - NVIS

NVIDIA • Town of Texas (WI)

On-site
USD 216,000 - 397,000
Equity
Senior AI Compute Engineer - NVIS
Senior AI Compute Engineer - NVIS

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 148,000 - 288,000