Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA

Westford (MA)

On-site

USD 152,000 - 241,500

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A leading technology company in Westford is seeking an individual to manage and optimize job scheduling systems in a high-performance computing environment. The ideal candidate will have a Bachelor’s in Computer Science and over five years of experience in supporting large-scale Linux-based infrastructures. Responsibilities include analyzing performance data, leading problem-solving efforts, and collaborating with customer teams. Competitive salaries and equity options are offered in addition to a comprehensive benefits package.

Qualifications

  • 5+ years of experience operating and supporting large-scale Linux-based compute infrastructure.
  • Hands-on experience supporting and tuning job scheduling systems in HPC environments.
  • Strong problem-solving skills and ability to analyze complex system behavior.

Responsibilities

  • Manage job scheduling systems in a multi-site environment.
  • Analyze infrastructure performance data to improve utilization.
  • Lead problem solving across scheduler and workload layers.

Skills

Job scheduling systems (LSF, Slurm)
Linux systems administration (CentOS/RHEL)
Problem-solving skills
Communication skills

Education

Bachelor’s degree in Computer Science or related field

Tools

Docker
Singularity
Podman

Job description

A leading technology company in Westford is seeking an individual to manage and optimize job scheduling systems in a high-performance computing environment. The ideal candidate will have a Bachelor’s in Computer Science and over five years of experience in supporting large-scale Linux-based infrastructures. Responsibilities include analyzing performance data, leading problem-solving efforts, and collaborating with customer teams. Competitive salaries and equity options are offered in addition to a comprehensive benefits package.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Durham (NC)

On-site
USD 152,000 - 241,500
Senior HPC Scheduler & Platform Reliability Engineer
Senior HPC Scheduler & Platform Reliability Engineer

NVIDIA • Austin (TX)

On-site
USD 152,000 - 241,500
Senior HPC Scheduler & Automation Engineer (LSF/Slurm)
Senior HPC Scheduler & Automation Engineer (LSF/Slurm)

NVIDIA Corporation • Austin (TX)

On-site
USD 152,000 - 287,500
Comprehensive benefits package
Equity opportunities
Systems Developer: HPC Scheduling & Platform Reliability
Systems Developer: HPC Scheduling & Platform Reliability

engineeringjobs.net, Inc. • New York (NY)

On-site
USD 110,000 - 170,000
Senior HPC Systems Engineer (SLURM & Red Hat)
Senior HPC Systems Engineer (SLURM & Red Hat)

Omega Enterprise Solutions, LLC • Corridor North (MD)

On-site
USD 90,000 - 120,000
HPC Systems Engineer - Scheduling & Performance
HPC Systems Engineer - Scheduling & Performance

Referment • New York (NY)

On-site
USD 90,000 - 130,000
Senior HPC Storage Engineer - Performance & Automation
Senior HPC Storage Engineer - Performance & Automation

NVIDIA • Austin (TX)

On-site
USD 184,000 - 287,500
Equity and comprehensive benefits package
Senior HPC SysAdmin | Slurm & Linux Expert (Remote)
Senior HPC SysAdmin | Slurm & Linux Expert (Remote)

Metasys Technologies • Boston (MA)

Remote
USD 79,556 - 113,307
Senior HPC Systems Integration Engineer
Senior HPC Systems Integration Engineer

Que Technology Group • Maryland

On-site
USD 120,000 - 160,000
Company Medical/Dental/Vision plans – Company paid
Short-term Disability, Long-term disability and Life Insurance – Company paid
Business/ First Class travel upgrade for 7 hour or longer flights
+5
Systems Developer: HPC, Kernel Tuning & Distributed Platforms
Systems Developer: HPC, Kernel Tuning & Distributed Platforms

engineeringjobs.net, Inc. • New York (NY)

On-site
USD 140,000 - 190,000