Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA

Durham (NC)

On-site

USD 152,000 - 241,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading technology company in Durham is seeking an experienced professional to optimize and scale job scheduling systems in a large-scale environment. You will manage job scheduling, identify performance bottlenecks, and implement process improvements. Ideal candidates will have a bachelor's degree in Computer Science, 5+ years of experience with Linux-based infrastructures, and strong problem-solving skills. The role offers a competitive salary and substantial benefits package.

Qualifications

  • 5+ years of experience operating large-scale Linux-based compute infrastructure.
  • Strong hands-on experience supporting job scheduling systems in HPC or silicon design environments.
  • Proficiency in Linux systems administration required.

Responsibilities

  • Manage, scale, and optimize job scheduling systems in a multi-site environment.
  • Analyze performance data to drive improvements in utilization and throughput.
  • Lead problem solving across scheduler, OS, and workload layers.

Skills

Job scheduling systems (LSF, Slurm)
Linux systems administration
Problem-solving skills
Communication skills
Data analysis

Education

Bachelor’s degree in Computer Science or related field

Tools

CentOS
Docker
Singularity
Podman

Job description

A leading technology company in Durham is seeking an experienced professional to optimize and scale job scheduling systems in a large-scale environment. You will manage job scheduling, identify performance bottlenecks, and implement process improvements. Ideal candidates will have a bachelor's degree in Computer Science, 5+ years of experience with Linux-based infrastructures, and strong problem-solving skills. The role offers a competitive salary and substantial benefits package.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Austin (TX)

On-site
USD 152,000 - 242,000
Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Westford (MA)

On-site
USD 152,000 - 242,000
Senior HPC Scheduler & Automation Engineer (LSF/Slurm)
Senior HPC Scheduler & Automation Engineer (LSF/Slurm)

NVIDIA Corporation • Austin (TX)

On-site
USD 152,000 - 288,000
Comprehensive benefits package
Equity opportunities
Senior HPC Engineer: Distributed Systems & Scale
Senior HPC Engineer: Distributed Systems & Scale

The Trade Desk, Inc. • Bellevue (WA)

On-site
USD 124,900 - 228,900
Comprehensive healthcare coverage
401k plan with matching
Paid time off up to 160 hours
+1
Senior Software Engineer, HPC & Cloud Infrastructure
Senior Software Engineer, HPC & Cloud Infrastructure

NVIDIA • Durham (NC)

On-site
USD 152,000 - 242,000
Senior HPC Scheduler & Automation Engineer - Equity
Senior HPC Scheduler & Automation Engineer - Equity

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior HPC Systems Engineer (SLURM & Red Hat)
Senior HPC Systems Engineer (SLURM & Red Hat)

Omega Enterprise Solutions, LLC • Corridor North (MD)

On-site
USD 90,000 - 120,000
HPC Scheduling Engineer: Kubernetes & Batch at Scale
HPC Scheduling Engineer: Kubernetes & Batch at Scale

NorthMark Strategies • Dallas (TX)

On-site
USD 90,000 - 130,000
Company-Paid Lunch Stipend
100% Employer-Paid Medical and Benefits
401(k) with company match
+2
Senior Data Center Scheduler – Large-Scale Infra
Senior Data Center Scheduler – Large-Scale Infra

Tract Capital Management, LP • Austin (TX)

Hybrid
USD 120,000 - 150,000
Senior Job Scheduling & Automation Engineer
Senior Job Scheduling & Automation Engineer

Jobsbridge • Scottsdale (AZ)

On-site