Senior HPC Scheduler (LSF/Slurm) Engineer

NVIDIA AI

Durham (NC)

On-site

USD 152,000 - 288,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Comprehensive benefits
Career growth opportunities

Job summary

NVIDIA is seeking an experienced engineer to optimize and operate a large-scale Linux compute infrastructure for EDA workloads. You will manage and scale job scheduling systems (LSF, Slurm) across multiple sites, analyze performance data, and drive reliability through automation and observability.

Ideal candidates have 5+ years of HPC experience, strong Linux skills, and the ability to communicate tradeoffs and metrics to engineering teams. This is a full-time role based in the United States.

Qualifications

  • Bachelor's degree in CS or related field or equivalent experience.
  • 5+ years operating large-scale Linux compute infrastructure.
  • Hands-on tuning of LSF, Slurm or similar schedulers in HPC or silicon design.
  • Proficient in Linux administration (CentOS/RHEL).
  • Strong analytical and communication skills to explain tradeoffs and reliability metrics.

Responsibilities

  • Manage, scale, and optimize job scheduling systems (LSF, Slurm) across sites.
  • Analyze scheduler and infra data to improve utilization and throughput.
  • Lead problem solving across scheduler, OS, and workload layers.
  • Automate recurring tasks to reduce manual effort and incidents.
  • Define and track metrics and SLOs for reliability with engineering stakeholders.
  • Contribute to standards, docs, and best practices across sites.
  • Collaborate with customer teams to clarify requirements and drive closure.

Skills

Linux administration
Scheduler tuning
HPC experience
Problem solving
Communication
OS troubleshooting

Education

Bachelor's degree in Computer Science or related field

Tools

LSF
Slurm
CentOS/RHEL

Job description

NVIDIA is seeking an experienced engineer to optimize and operate a large-scale Linux compute infrastructure for EDA workloads. You will manage and scale job scheduling systems (LSF, Slurm) across multiple sites, analyze performance data, and drive reliability through automation and observability.

Ideal candidates have 5+ years of HPC experience, strong Linux skills, and the ability to communicate tradeoffs and metrics to engineering teams. This is a full-time role based in the United States.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Scheduler Engineer (LSF/Slurm) - Hybrid & Equity
Senior HPC Scheduler Engineer (LSF/Slurm) - Hybrid & Equity

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Equity
Benefits package
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 241,500
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Austin (TX)

On-site
USD 152,000 - 242,000
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Westford (MA)

On-site
USD 152,000 - 242,000
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Durham (NC)

On-site
USD 152,000 - 242,000
Senior HPC Scheduler & Reliability Engineer — Equity Options
Senior HPC Scheduler & Reliability Engineer — Equity Options

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 241,500
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Equity
Benefits package
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA AI • Durham (NC)

On-site
USD 152,000 - 288,000
Equity
Comprehensive benefits
Career growth opportunities
Senior LSF Scheduler Architect - HPC & Multi-Cluster
Senior LSF Scheduler Architect - HPC & Multi-Cluster

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Hybrid work model
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA Corporation • Austin (TX)

On-site
USD 152,000 - 288,000
Comprehensive benefits package
Equity opportunities