Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)

Rune

Mountain View (CA)

Hybrid

USD 175,000 - 260,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Rune is seeking a hands-on engineer to own customer-facing technical support for our GPU clusters, covering networking, provisioning, and GPU health. You will diagnose issues, resolve them, and avoid hand-offs by maintaining technical judgment.

You will work on RoCEv2, NCCL, and NCCL benchmarking, while managing driver/firmware compatibility across a fleet. This is a hybrid/remote role with a focus on delivering reliable compute at solar sites.

Qualifications

  • Hands-on experience with RDMA fabrics, RoCEv2, ECN/DCQCN, PFC, and DCBX configuration.
  • Experience debugging multi-node GPU interconnect performance with NCCL and RDMA benchmarking tools.
  • Knowledge of GPU interconnect topology and NUMA/PCIe affinity.
  • Experience reading GPU health telemetry and managing driver/firmware across a fleet.
  • Strong Linux systems administration and fleet automation (Ansible or equivalent).

Responsibilities

  • Serve as the primary technical responder for customer-reported networking, GPU-interconnect, and node issues on live clusters.
  • Diagnose multi-node NCCL and RDMA performance regressions and distinguish hypotheses from confirmed findings.
  • Configure and validate RoCEv2 congestion control and reconcile NIC settings with switch fabric behavior.
  • Diagnose GPU hardware faults using DCGM telemetry and Xid codes, differentiating hardware failures from software issues.
  • Own driver and firmware compatibility across the fleet and validate updates.
  • Provision, re-provision, and decommission GPU nodes with validation testing.
  • Triage edge-network and jumpbox access issues affecting provisioning workflows.
  • Write incident updates and post-resolution summaries for customer teams.

Skills

RoCEv2
NCCL
NVIDIA DCGM
Linux admin
Ansible

Tools

nvidia-smi
SSH

Job description

Rune is seeking a hands-on engineer to own customer-facing technical support for our GPU clusters, covering networking, provisioning, and GPU health. You will diagnose issues, resolve them, and avoid hand-offs by maintaining technical judgment.

You will work on RoCEv2, NCCL, and NCCL benchmarking, while managing driver/firmware compatibility across a fleet. This is a hybrid/remote role with a focus on delivering reliable compute at solar sites.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infrastructure Engineer
GPU Infrastructure Engineer

Rune • Mountain View (CA)

Hybrid
USD 175,000 - 260,000
Senior GPU Infra Engineer — Remote
Senior GPU Infra Engineer — Remote

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Senior GPU Infrastructure Support Engineer
Senior GPU Infrastructure Support Engineer

Nscale • San Francisco (CA)

On-site
USD 120,000 - 170,000
Equity
Remote-friendly team
Flexible workplace
Remote Infra Operations Engineer - GPU Cloud
Remote Infra Operations Engineer - GPU Cloud

Nscale • Barstow (TX)

Remote
USD 80,000 - 110,000
Competitive pay with equity
Flexible workplace
Dynamic progression plan
Senior GPU Compute Fleet Engineer | Remote-First
Senior GPU Compute Fleet Engineer | Remote-First

Boundless Networks • United States

Remote
USD 175,000 - 250,000
Equity
Health, dental, vision
Flexible PTO
+2
Remote Senior GPU Network Engineer – Fabric & Cluster Builds
Remote Senior GPU Network Engineer – Fabric & Cluster Builds

REALM • United States

On-site
USD 170,000 - 230,000
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale

Luma • Redwood City (CA)

Remote
USD 180,000 - 240,000
Remote NOC Engineer - GPU Clusters & Incident Response
Remote NOC Engineer - GPU Clusters & Incident Response

REALM • United States

On-site
USD 70,000 - 110,000
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior GPU Cluster Architect (Remote)
Senior GPU Cluster Architect (Remote)

Orion Placement • United States

On-site
USD 150,000 - 230,000
Remote nationwide work
Bonus potential
Equity opportunity