Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA

Austin (TX)

On-site

USD 152,000 - 241,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology company seeks an experienced candidate to manage and optimize job scheduling systems in a large-scale environment. This role involves analyzing performance data, resolving service-impacting issues, and implementing automation. Ideal candidates will have a Bachelor's degree in Computer Science and over 5 years of experience with Linux-based infrastructures and scheduling systems like LSF and Slurm. NVIDIA offers competitive salaries, equity, and comprehensive benefits.

Qualifications

  • 5+ years of experience operating large-scale Linux-based compute infrastructure.
  • Experience with job scheduling systems (LSF, Slurm) in HPC or silicon design.

Responsibilities

  • Manage job scheduling systems in a multi-site environment.
  • Analyze performance data to improve utilization and throughput.
  • Implement automation to reduce manual effort.

Skills

Linux systems administration
Job scheduling systems (LSF, Slurm)
Problem-solving skills
Communication skills

Education

Bachelor’s degree in Computer Science or related field

Tools

CentOS/RHEL
Docker
Slurm
LSF

Job description

A technology company seeks an experienced candidate to manage and optimize job scheduling systems in a large-scale environment. This role involves analyzing performance data, resolving service-impacting issues, and implementing automation. Ideal candidates will have a Bachelor's degree in Computer Science and over 5 years of experience with Linux-based infrastructures and scheduling systems like LSF and Slurm. NVIDIA offers competitive salaries, equity, and comprehensive benefits.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Durham (NC)

On-site
USD 152,000 - 242,000
Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Westford (MA)

On-site
USD 152,000 - 242,000
Senior HPC Scheduler & Automation Engineer (LSF/Slurm)
Senior HPC Scheduler & Automation Engineer (LSF/Slurm)

NVIDIA Corporation • Austin (TX)

On-site
USD 152,000 - 288,000
Comprehensive benefits package
Equity opportunities
Senior HPC Scheduler & Automation Engineer - Equity
Senior HPC Scheduler & Automation Engineer - Equity

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Durham (NC)

On-site
USD 152,000 - 242,000
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Austin (TX)

On-site
USD 152,000 - 242,000
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Westford (MA)

On-site
USD 152,000 - 242,000
Senior GPU HPC Scheduler Engineer | Equity Eligible
Senior GPU HPC Scheduler Engineer | Equity Eligible

NVIDIA • Redmond (WA)

On-site
USD 152,000 - 242,000
Senior HPC Systems Engineer (SLURM & Red Hat)
Senior HPC Systems Engineer (SLURM & Red Hat)

Omega Enterprise Solutions, LLC • Corridor North (MD)

On-site
USD 90,000 - 120,000