Senior AMD GPU Infra Engineer for ROCm AI Stack

Evergrid

New York (NY)

Hybrid

USD 180,000 - 220,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
Paid time off
Home office stipend

Job summary

Evergrid seeks an experienced HPC GPU operations engineer to own the health, reliability, and performance of AMD GPU clusters. You will lead Linux systems, kernel debugging, and ROCm-based ML stack across production-scale AI workloads.

You will be on-call for outages, optimize GPU topology, and work with data centers and vendors. A strong background in HPC, Python, Bash, and ROCm is required.

Qualifications

  • 5+ years in HPC or GPU cluster operations.
  • BS/MS in CS/CE/EE or related field.
  • Hands-on with AMD MI-series GPUs and driver/kernel debugging.
  • Strong Linux internals, kernel modules, and performance tuning.
  • Experience securing production infra: VPNs, firewalls, SSH, identity systems.
  • Proficient in Bash and Python for automation.
  • Familiarity with ROCm stack (PyTorch ROCm, JAX ROCm, RCCL, MIOpen).
  • Experience debugging high-performance networking and RDMA.

Responsibilities

  • Own system health and on-call response for outages and GPU failures.
  • Be the primary point of contact for POC and active customers' GPU clusters.
  • Diagnose and resolve issues to minimize downtime and uphold SLA.
  • Monitor GPU health, thermals, PCIe topology, memory errors, cluster load.
  • Coordinate repairs with data centers and hardware vendors.

Skills

HPC GPU clusters
Linux systems engineering
Python automation
Bash scripting
ROCm / ROCm stack
AMD MI-series GPUs
Kernel debugging
On-call incident response

Education

Bachelor's or Master's in CS/CE/EE

Tools

InfiniBand/RoCE
SSH/VPNs

Job description

Evergrid seeks an experienced HPC GPU operations engineer to own the health, reliability, and performance of AMD GPU clusters. You will lead Linux systems, kernel debugging, and ROCm-based ML stack across production-scale AI workloads.

You will be on-call for outages, optimize GPU topology, and work with data centers and vendors. A strong background in HPC, Python, Bash, and ROCm is required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU AI/ML Vision & Quality Engineer
Senior GPU AI/ML Vision & Quality Engineer

AMD • Santa Clara (CA)

On-site
USD 140,000 - 210,000
Senior HPC & AMD Infrastructure Engineer
Senior HPC & AMD Infrastructure Engineer

Evergrid • New York (NY)

Hybrid
USD 180,000 - 220,000
Health insurance
Paid time off
Home office stipend
Staff GPU & AI/ML Software Engineer for CV, QA & ROCm
Staff GPU & AI/ML Software Engineer for CV, QA & ROCm

Advanced Micro Devices, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 120,000 - 190,000
ROCm GPU Libraries Fellow: AI & HPC Architect Lead
ROCm GPU Libraries Fellow: AI & HPC Architect Lead

Advanced Micro Devices, Inc. • Oregon

Hybrid
USD 200,000 - 320,000
Senior GPU/AI Systems Engineer - Performance & ML
Senior GPU/AI Systems Engineer - Performance & ML

AMD • Santa Clara (CA)

On-site
USD 170,000 - 250,000
AMD Benefits
Fellow, ROCm GPU Libraries — AI/HPC Architect
Fellow, ROCm GPU Libraries — AI/HPC Architect

AMD • Oregon

On-site
USD 180,000 - 260,000
Remote Senior Product Manager, ROCm AI Inference
Remote Senior Product Manager, ROCm AI Inference

AMD • Santa Clara (CA)

Hybrid
USD 120,000 - 150,000
Senior GPU Software Engineer — HPC Libraries
Senior GPU Software Engineer — HPC Libraries

AMD • San Jose (CA)

Hybrid
USD 180,000 - 250,000
Senior GPU AI Software Engineer — Kernels to Scale AI
Senior GPU AI Software Engineer — Kernels to Scale AI

Advanced Micro Devices, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Senior Data Center GPU Validation & Debug Engineer
Senior Data Center GPU Validation & Debug Engineer

Advanced Micro Devices • Austin (TX)

Hybrid
USD 120,000 - 180,000