HPC Operations Engineer - Optimize Compute Systems (Equity)

NVIDIA

United States

On-site

USD 124,000 - 242,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Comprehensive benefits

Job summary

NVIDIA is seeking an HPC Operations Engineer to ensure flawless operation of a high-performance computing environment. Join a world-class team powering semiconductor build and advanced engineering workflows, contributing to innovative computing developments.

You will provide first-line support for HPC users, troubleshoot failures, monitor health and queues, and maintain runbooks and guidelines. The role emphasizes collaboration, process adherence, and opportunities to influence workflow

Qualifications

  • Bachelor's degree in CS/IT/Engineering or equivalent experience.
  • 2+ years supporting Linux-based production environments.
  • Solid Linux system administration fundamentals (RHEL/CentOS and/or Ubuntu).
  • Ability to troubleshoot technical issues methodically and escalate as needed.
  • Experience interacting directly with users in a technical support or operations role.
  • Strong written communication and ability to produce clear documentation.
  • Follow established processes with attention to detail.

Responsibilities

  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, resource constraints, and performance concerns.
  • Perform triage of infrastructure incidents and engage SMEs when needed.
  • Monitor system health, queues, node status, and service availability.
  • Complete maintenance, patching, and configuration updates per procedures.
  • Develop and maintain runbooks, knowledge base, and guidelines.
  • Identify recurring issues and propose workflow improvements.

Skills

Linux administration
Customer support
Documentation
Scripting basics

Education

Bachelor's degree or equivalent

Tools

LSF
Slurm
NFS
LDAP
Bash
Python

Job description

NVIDIA is seeking an HPC Operations Engineer to ensure flawless operation of a high-performance computing environment. Join a world-class team powering semiconductor build and advanced engineering workflows, contributing to innovative computing developments.

You will provide first-line support for HPC users, troubleshoot failures, monitor health and queues, and maintain runbooks and guidelines. The role emphasizes collaboration, process adherence, and opportunities to influence workflow

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Performance Engineer: Optimize Compilers for GPUs (Equity)
HPC Performance Engineer: Optimize Compilers for GPUs (Equity)

NVIDIA Gruppe • Oregon (WI)

Hybrid
USD 152,000 - 242,000
Senior HPC Middleware Engineer — Equity
Senior HPC Middleware Engineer — Equity

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
HPC Operations Engineer
HPC Operations Engineer

NVIDIA • United States

On-site
USD 124,000 - 242,000
Equity
Comprehensive benefits
Senior HPC Scheduler & Reliability Engineer — Equity Options
Senior HPC Scheduler & Reliability Engineer — Equity Options

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 241,500
HPC Deployment & AI Systems Lead — Equity Options
HPC Deployment & AI Systems Lead — Equity Options

NVIDIA • Virginia (MN)

On-site
USD 216,000 - 397,000
Equity
Benefits
Senior Linux Systems Engineer — HPC & Automation (Equity)
Senior Linux Systems Engineer — HPC & Automation (Equity)

NVIDIA • Washington

On-site
USD 235,000 - 357,000
Equity
Benefits package
HPC Performance Engineer: Compiler & GPU Optimization
HPC Performance Engineer: Compiler & GPU Optimization

NVIDIA • Town of Texas (WI)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior HPC Middleware Engineer – Equity & Performance
Senior HPC Middleware Engineer – Equity & Performance

NVIDIA • Illinois

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior HPC Scheduler Engineer (LSF/Slurm) - Hybrid & Equity
Senior HPC Scheduler Engineer (LSF/Slurm) - Hybrid & Equity

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Equity
Benefits package
Senior HPC Architect - GPU Compute, Equity Eligible
Senior HPC Architect - GPU Compute, Equity Eligible

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Inclusive work environment
Comprehensive benefits