Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA

Austin (TX)

On-site

USD 152,000 - 241,500

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

A technology company seeks an experienced candidate to manage and optimize job scheduling systems in a large-scale environment. This role involves analyzing performance data, resolving service-impacting issues, and implementing automation. Ideal candidates will have a Bachelor's degree in Computer Science and over 5 years of experience with Linux-based infrastructures and scheduling systems like LSF and Slurm. NVIDIA offers competitive salaries, equity, and comprehensive benefits.

Qualifications

  • 5+ years of experience operating large-scale Linux-based compute infrastructure.
  • Experience with job scheduling systems (LSF, Slurm) in HPC or silicon design.

Responsibilities

  • Manage job scheduling systems in a multi-site environment.
  • Analyze performance data to improve utilization and throughput.
  • Implement automation to reduce manual effort.

Skills

Linux systems administration
Job scheduling systems (LSF, Slurm)
Problem-solving skills
Communication skills

Education

Bachelor’s degree in Computer Science or related field

Tools

CentOS/RHEL
Docker
Slurm
LSF

Job description

A technology company seeks an experienced candidate to manage and optimize job scheduling systems in a large-scale environment. This role involves analyzing performance data, resolving service-impacting issues, and implementing automation. Ideal candidates will have a Bachelor's degree in Computer Science and over 5 years of experience with Linux-based infrastructures and scheduling systems like LSF and Slurm. NVIDIA offers competitive salaries, equity, and comprehensive benefits.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Durham (NC)

On-site
USD 152,000 - 241,500
Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Westford (MA)

On-site
USD 152,000 - 241,500
Senior HPC Scheduler & Automation Engineer (LSF/Slurm)
Senior HPC Scheduler & Automation Engineer (LSF/Slurm)

NVIDIA Corporation • Austin (TX)

On-site
USD 152,000 - 287,500
Comprehensive benefits package
Equity opportunities
Senior HPC Scheduler Engineer (LSF/Slurm) - Hybrid & Equity
Senior HPC Scheduler Engineer (LSF/Slurm) - Hybrid & Equity

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Equity
Benefits package
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Durham (NC)

On-site
USD 152,000 - 241,500
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Westford (MA)

On-site
USD 152,000 - 241,500
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Austin (TX)

On-site
USD 152,000 - 241,500
Systems Developer: HPC Scheduling & Platform Reliability
Systems Developer: HPC Scheduling & Platform Reliability

engineeringjobs.net, Inc. • New York (NY)

On-site
USD 110,000 - 170,000
Senior GPU HPC Scheduler Engineer | Equity Eligible
Senior GPU HPC Scheduler Engineer | Equity Eligible

NVIDIA • Redmond (WA)

On-site
USD 152,000 - 241,500
Equity
Benefits
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

On-site
USD 152,000 - 288,000
Equity
Benefits package