Senior HPC Scheduler & Reliability Engineer — Equity Options

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 152,000 - 241,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Gruppe in Santa Clara is hiring for a role in their Hardware Infrastructure EDA Compute team to optimize workload scheduling systems and improve overall service reliability. The successful candidate will manage and scale job scheduling systems while driving measurable improvements in efficiency through automation and metric tracking.

A Bachelor’s degree in Computer Science or related field and 5+ years of relevant experience are required. Knowledge of LSF and Slurm scheduling systems is vital for this position.

This role offers competitive compensation and a comprehensive benefits package.

Qualifications

  • Minimum 5+ years of experience operating and supporting large-scale Linux-based compute infrastructure.
  • Strong hands-on experience supporting and tuning job scheduling systems (LSF, Slurm) in HPC or silicon design environments.
  • Clear and effective communication skills to articulate technical tradeoffs.

Responsibilities

  • Manage, scale, and optimize job scheduling systems in a multi-site environment.
  • Analyze scheduler and infrastructure performance data to drive improvements.
  • Identify operational challenges and implement automation to reduce manual effort.

Skills

Linux systems administration
Job scheduling systems (LSF, Slurm)
Problem-solving skills
Monitoring pipelines
Observability systems

Education

Bachelor’s degree in Computer Science or related field

Tools

Docker
CentOS/RHEL

Job description

NVIDIA Gruppe in Santa Clara is hiring for a role in their Hardware Infrastructure EDA Compute team to optimize workload scheduling systems and improve overall service reliability. The successful candidate will manage and scale job scheduling systems while driving measurable improvements in efficiency through automation and metric tracking.

A Bachelor’s degree in Computer Science or related field and 5+ years of relevant experience are required. Knowledge of LSF and Slurm scheduling systems is vital for this position.

This role offers competitive compensation and a comprehensive benefits package.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 241,500
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Austin (TX)

On-site
USD 152,000 - 242,000
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Durham (NC)

On-site
USD 152,000 - 242,000
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA • Westford (MA)

On-site
USD 152,000 - 242,000
Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Austin (TX)

On-site
USD 152,000 - 242,000
Senior GPU HPC Scheduler Engineer | Equity Eligible
Senior GPU HPC Scheduler Engineer | Equity Eligible

NVIDIA • Redmond (WA)

On-site
USD 152,000 - 242,000
Senior HPC and LSF Operations Engineer
Senior HPC and LSF Operations Engineer

NVIDIA Corporation • Austin (TX)

On-site
USD 152,000 - 288,000
Comprehensive benefits package
Equity opportunities
Senior HPC-AI Systems Architect (Equity)
Senior HPC-AI Systems Architect (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Senior HPC Cluster Engineer
Senior HPC Cluster Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior HPC Middleware Engineer — Equity
Senior HPC Middleware Engineer — Equity

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000