Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA

Durham (NC)

On-site

USD 152,000 - 241,500

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

A leading technology company in Durham is seeking an experienced professional to optimize and scale job scheduling systems in a large-scale environment. You will manage job scheduling, identify performance bottlenecks, and implement process improvements. Ideal candidates will have a bachelor's degree in Computer Science, 5+ years of experience with Linux-based infrastructures, and strong problem-solving skills. The role offers a competitive salary and substantial benefits package.

Qualifications

  • 5+ years of experience operating large-scale Linux-based compute infrastructure.
  • Strong hands-on experience supporting job scheduling systems in HPC or silicon design environments.
  • Proficiency in Linux systems administration required.

Responsibilities

  • Manage, scale, and optimize job scheduling systems in a multi-site environment.
  • Analyze performance data to drive improvements in utilization and throughput.
  • Lead problem solving across scheduler, OS, and workload layers.

Skills

Job scheduling systems (LSF, Slurm)
Linux systems administration
Problem-solving skills
Communication skills
Data analysis

Education

Bachelor’s degree in Computer Science or related field

Tools

CentOS
Docker
Singularity
Podman

Job description

A leading technology company in Durham is seeking an experienced professional to optimize and scale job scheduling systems in a large-scale environment. You will manage job scheduling, identify performance bottlenecks, and implement process improvements. Ideal candidates will have a bachelor's degree in Computer Science, 5+ years of experience with Linux-based infrastructures, and strong problem-solving skills. The role offers a competitive salary and substantial benefits package.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Austin (TX)

On-site
USD 152,000 - 241,500
Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Westford (MA)

On-site
USD 152,000 - 241,500
Senior HPC Scheduler & Automation Engineer (LSF/Slurm)
Senior HPC Scheduler & Automation Engineer (LSF/Slurm)

NVIDIA Corporation • Austin (TX)

On-site
USD 152,000 - 287,500
Comprehensive benefits package
Equity opportunities
Systems Developer: HPC Scheduling & Platform Reliability
Systems Developer: HPC Scheduling & Platform Reliability

engineeringjobs.net, Inc. • New York (NY)

On-site
USD 110,000 - 170,000
HPC Systems Engineer - Scheduling & Performance
HPC Systems Engineer - Scheduling & Performance

Referment • New York (NY)

On-site
USD 90,000 - 130,000
Senior HPC Scheduler Engineer (LSF/Slurm) - Hybrid & Equity
Senior HPC Scheduler Engineer (LSF/Slurm) - Hybrid & Equity

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Equity
Benefits package
Senior HPC Systems Engineer (SLURM & Red Hat)
Senior HPC Systems Engineer (SLURM & Red Hat)

Omega Enterprise Solutions, LLC • Corridor North (MD)

On-site
USD 90,000 - 120,000
Systems Developer: HPC, Kernel Tuning & Distributed Platforms
Systems Developer: HPC, Kernel Tuning & Distributed Platforms

engineeringjobs.net, Inc. • New York (NY)

On-site
USD 140,000 - 190,000
Systems Developer: HPC Workload & Fleet Management
Systems Developer: HPC Workload & Fleet Management

engineeringjobs.net, Inc. • New York (NY)

On-site
USD 120,000 - 160,000
Senior Data Center Scheduler – Large-Scale Infra
Senior Data Center Scheduler – Large-Scale Infra

Tract Capital Management, LP • Austin (TX)

Hybrid
USD 120,000 - 150,000
100% employer-covered medical, dental, and vision insurance
401K program
Unlimited PTO