AI Compute Infrastructure Engineer

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 124,000 - 196,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking an AI Compute Engineer to join its Infrastructure Specialists team. The role involves deploying, managing, and validating AI Compute/HPC infrastructure in Linux environments for new and existing customers, while coordinating with customers, partners, and internal teams.

You will act as the domain expert during planning and implementation, perform knowledge transfers, and provide feedback to engineering teams.

Qualifications

  • 4+ years providing in-depth support and deployment services for hardware and software products.
  • Linux system administration, process management, package management, task scheduling, kernel management, boot procedures/troubleshooting, performance reporting/logging, and advanced networking.
  • Cluster management and bare-metal provisioning technologies; BCM is a bonus.
  • Four-year degree from an accredited university or college in Computer Science, Electrical or Computer Engineering or equivalent experience.
  • Scripting proficiency (Bash, Python, Ansible, etc.).
  • Excellent interpersonal and organizational skills with ability to work with limited supervision.
  • Experience with schedulers such as SLURM, LSF, UGE.
  • Willingness to travel to customer sites within the United States up to 20%.
  • Experience with benchmarking tools (HPL, NCCL tests, MLPerf) and Kubernetes.

Responsibilities

  • Deploying, managing, and validating AI Compute/HPC infrastructure in Linux-based environments for new and existing customers.
  • Be the domain expert with customers during planning calls through implementation.
  • Handover-related documentation and perform knowledge transfers required to support customers as they begin rolling out some of the most sophisticated systems in the world.
  • Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements.

Skills

Linux administration
Scripting
Cluster management
Networking & troubleshooting
Documentation & knowledge transfer
Interpersonal skills
Organizational skills
Travel up to 20%

Education

Bachelor's degree in CS/EE or equivalent

Tools

SLURM
LSF
UGE
BCM Base Command Manager
HPL
NCCL tests
MLPerf
Kubernetes
InfiniBand
GPFS/Lustre

Job description

NVIDIA is seeking an AI Compute Engineer to join its Infrastructure Specialists team. The role involves deploying, managing, and validating AI Compute/HPC infrastructure in Linux environments for new and existing customers, while coordinating with customers, partners, and internal teams.

You will act as the domain expert during planning and implementation, perform knowledge transfers, and provide feedback to engineering teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Compute Systems Engineer – HPC & Linux Infra
AI Compute Systems Engineer – HPC & Linux Infra

NVIDIA • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Senior AI Compute Engineer — AI Infrastructure Lead
Senior AI Compute Engineer — AI Infrastructure Lead

NVIDIA • United States

Remote
USD 148,000 - 236,000
Senior AI Compute Deployment Engineer
Senior AI Compute Deployment Engineer

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity and Benefits
AI Infrastructure Engineer — GPU Compute Lead
AI Infrastructure Engineer — GPU Compute Lead

NVIDIA Corporation • Town of Italy (NY)

Hybrid
USD 74,000 - 168,000
Senior AI Compute Field Engineer
Senior AI Compute Field Engineer

NVIDIA • California (MO)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

NVIDIA • United States

Remote
USD 120,000 - 180,000
Senior AI Compute Engineer — Lead Large-Scale AI Infra
Senior AI Compute Engineer — Lead Large-Scale AI Infra

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 148,000 - 288,000
Equity
Benefits
Senior AI Compute Engineer - Field Deployments & HPC
Senior AI Compute Engineer - Field Deployments & HPC

NVIDIA • Georgia

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Compute Engineer - Field Deployment Lead
Senior AI Compute Engineer - Field Deployment Lead

NVIDIA • New Jersey

On-site
USD 184,000 - 357,000
Equity
Benefits
AI Compute Engineer - NVIS
AI Compute Engineer - NVIS

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Equity
Benefits