Senior HPC & AMD Infrastructure Engineer

Evergrid

New York (NY)

Ibrido

USD 180.000 - 220.000

Tempo pieno

10 giorni fa
Generatore di candidature

Una candidatura completa in un minuto — curriculum e lettera di presentazione personalizzati, pronti da inviare.

Supera i filtri ATS

Vantaggi offerti da questo lavoro

Health insurance
Paid time off
Home office stipend

Descrizione del lavoro

Evergrid seeks an experienced HPC GPU operations engineer to own the health, reliability, and performance of AMD GPU clusters. You will lead Linux systems, kernel debugging, and ROCm-based ML stack across production-scale AI workloads.

You will be on-call for outages, optimize GPU topology, and work with data centers and vendors. A strong background in HPC, Python, Bash, and ROCm is required.

Competenze

  • 5+ years in HPC or GPU cluster operations.
  • BS/MS in CS/CE/EE or related field.
  • Hands-on with AMD MI-series GPUs and driver/kernel debugging.
  • Strong Linux internals, kernel modules, and performance tuning.
  • Experience securing production infra: VPNs, firewalls, SSH, identity systems.
  • Proficient in Bash and Python for automation.
  • Familiarity with ROCm stack (PyTorch ROCm, JAX ROCm, RCCL, MIOpen).
  • Experience debugging high-performance networking and RDMA.

Mansioni

  • Own system health and on-call response for outages and GPU failures.
  • Be the primary point of contact for POC and active customers' GPU clusters.
  • Diagnose and resolve issues to minimize downtime and uphold SLA.
  • Monitor GPU health, thermals, PCIe topology, memory errors, cluster load.
  • Coordinate repairs with data centers and hardware vendors.

Conoscenze

HPC GPU clusters
Linux systems engineering
Python automation
Bash scripting
ROCm / ROCm stack
AMD MI-series GPUs
Kernel debugging
On-call incident response

Formazione

Bachelor's or Master's in CS/CE/EE

Strumenti

InfiniBand/RoCE
SSH/VPNs

Descrizione del lavoro

United States | Hybrid or Remote | Full-time

The role

You'll own the health, reliability, and performance of Evergrid's AMD-only GPU compute clusters.

You're the primary custodian of our high-density accelerator environments. The work spans hardware operations, Linux systems engineering, distributed infrastructure, and ML workloads. It covers GPU bring-up and kernel-level debugging, as well as maintaining and optimizing the ROCm-based ML stack behind production-scale AI. If you like getting maximum performance out of hardware, debugging GPUs at scale, and shipping world-class AI infrastructure, this role is for you.

What you'll own
System health and reliability (SRE)
  • Primary on-call response for outages, GPU failures, node crashes, and cluster-wide incidents.

  • Being the key point person for POC and active customers' GPU clusters.

  • Fast diagnosis and resolution that minimizes downtime and keeps SLA-level reliability.

  • Monitoring for GPU health, thermals, PCIe topology, memory errors, and cluster load.

  • Repairs, RMAs, and physical maintenance, coordinated with data center operators, hardware vendors, and on-site technicians.

Linux and network administration
  • Installing, patching, and maintaining Linux (Ubuntu, CentOS, RHEL) across large GPU node fleets.

  • Kernel tuning, consistent OS configuration, and fleet automation at scale.

  • Secure networking: VPNs, firewalls (iptables/firewalld), SSH hardening, and routing.

  • Identity and access systems (LDAP, FreeIPA, Active Directory).

  • Distributed storage (NFS, GPFS, Lustre).

AMD GPU and ML stack engineering (ROCm-first)
  • Deployment and bring-up of new GPU nodes, including BIOS configuration, NUMA tuning, and topology validation.

  • AMD GPU drivers, kernel modules, and the ROCm runtime across production fleets.

  • The AMD ML stack: ROCm, PyTorch (ROCm builds), JAX (ROCm/XLA), RCCL, hipBLAS/hipDNN, MIOpen, and supporting runtimes.

  • Debugging complex failures across GPUs, compilers, ML frameworks, and distributed training and inference. Examples include RCCL hangs, HIP memory faults, ROCm kernel crashes, framework build and link issues, and vLLM build failures on ROCm.

  • Infrastructure that supports both research iteration and production reliability, built with the ML and platform teams.

What you bring
  • 5+ years in HPC, GPU cluster operations, Linux systems engineering, or similar roles.

  • A bachelor's or master's in Computer Science, Computer Engineering, Electrical Engineering, or a related field.

  • Deep hands-on experience with AMD MI-series GPUs, including driver and kernel-level debugging.

  • Strong grasp of Linux internals, kernel modules, hardware bring-up, and performance tuning.

  • Experience securing and operating production infrastructure: VPNs, firewalls, SSH, and identity systems.

  • Proficiency in Bash and Python for automation, tooling, and operations.

  • Strong familiarity with ML stacks and runtime behavior in ROCm environments (ROCm/HIP, MIOpen, RCCL, PyTorch, JAX).

  • Experience debugging high-performance networking and RDMA (InfiniBand or RoCE), including cluster-level communication failures that affect distributed training.

Bonus
  • Schedulers and orchestration (Slurm, Kubernetes).

  • Model serving and inference optimization on ROCm (vLLM, SGLang).

  • Configuration management and IaC (Ansible, SaltStack, Terraform).

  • Supporting ML research or production AI teams at a startup or high-growth company.

Benefits & Compensation
  • Compensation: $180,000 - $220,000

  • Health insurance

  • Paid time off and paid holidays

  • Home office stipend

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior HPC & GPU Infrastructure Engineer
Senior HPC & GPU Infrastructure Engineer

Sciforium • San Francisco (CA)

In loco
USD 180.000 - 240.000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior AMD GPU Infra Engineer for ROCm AI Stack
Senior AMD GPU Infra Engineer for ROCm AI Stack

Evergrid • New York (NY)

Ibrido
USD 180.000 - 220.000
Health insurance
Paid time off
Home office stipend
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

In loco
USD 150.000 - 220.000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Hyperbolic • San Francisco (CA)

In loco
USD 180.000 - 260.000
Staff Software Engineer, GPU Inference
Staff Software Engineer, GPU Inference

Cerebras • Sunnyvale (CA)

In loco
USD 180.000 - 280.000
Senior Staff Software Development Engineer- GPU/AI/ML
Senior Staff Software Development Engineer- GPU/AI/ML

Advanced Micro Devices, Inc. • Santa Clara (CA)

In loco
USD 180.000 - 240.000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect AI • San Francisco (CA)

In loco
USD 150.000 - 300.000
Principal / Senior GPU SW Performance Engineer — Post‑Training
Principal / Senior GPU SW Performance Engineer — Post‑Training

AMD • San Jose (CA)

In loco
USD 120.000 - 160.000
Senior GPU Software Performance Engineer – Post-Training
Senior GPU Software Performance Engineer – Post-Training

AMD • San Jose (CA)

In loco
USD 170.000 - 250.000
Senior GPU Performance / Kernel Engineer
Senior GPU Performance / Kernel Engineer

Designworks Talent LLC • Bellevue (KY)

Ibrido
USD 170.000 - 250.000
Hybrid work model
Medical, dental, vision insurance