GPU Infrastructure Engineer

Rune

Mountain View (CA)

Hybrid

USD 175,000 - 260,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Rune is seeking a hands-on engineer to own customer-facing technical support for our GPU clusters, covering networking, provisioning, and GPU health. You will diagnose issues, resolve them, and avoid hand-offs by maintaining technical judgment.

You will work on RoCEv2, NCCL, and NCCL benchmarking, while managing driver/firmware compatibility across a fleet. This is a hybrid/remote role with a focus on delivering reliable compute at solar sites.

Qualifications

  • Hands-on experience with RDMA fabrics, RoCEv2, ECN/DCQCN, PFC, and DCBX configuration.
  • Experience debugging multi-node GPU interconnect performance with NCCL and RDMA benchmarking tools.
  • Knowledge of GPU interconnect topology and NUMA/PCIe affinity.
  • Experience reading GPU health telemetry and managing driver/firmware across a fleet.
  • Strong Linux systems administration and fleet automation (Ansible or equivalent).

Responsibilities

  • Serve as the primary technical responder for customer-reported networking, GPU-interconnect, and node issues on live clusters.
  • Diagnose multi-node NCCL and RDMA performance regressions and distinguish hypotheses from confirmed findings.
  • Configure and validate RoCEv2 congestion control and reconcile NIC settings with switch fabric behavior.
  • Diagnose GPU hardware faults using DCGM telemetry and Xid codes, differentiating hardware failures from software issues.
  • Own driver and firmware compatibility across the fleet and validate updates.
  • Provision, re-provision, and decommission GPU nodes with validation testing.
  • Triage edge-network and jumpbox access issues affecting provisioning workflows.
  • Write incident updates and post-resolution summaries for customer teams.

Skills

RoCEv2
NCCL
NVIDIA DCGM
Linux admin
Ansible

Tools

nvidia-smi
SSH

Job description

Location: Mountain View, CA, Hybrid, or Remote

About Rune

Every solar and wind power plant is a latent data center. Rune connects wasted power to the compute that needs it. Solar plants generate electricity the grid can't take: clipped by inverters, curtailed by operators, gone before it ever reaches a meter. Meanwhile, AI needs more power than new construction can deliver in time.

Rune deploys hardware, software, and financial solutions to put this power to work. We build modular micro data centers that install directly at solar sites, bypassing the traditional grid entirely and turning stranded power into compute. We're backed by Spark Capital, Union Square Ventures, and Lowercarbon Capital.

The Role

We're looking for a hands-on engineer to own customer-facing technical support for our GPU clusters: networking, node provisioning, and GPU health. You'll be the primary responder to ensure our customers’ clusters are performing as expected, including diagnosing issues, resolving them, or routing them correctly, without needing to hand off the technical judgment to someone else.

Key Responsibilities
  • Serve as the primary technical responder for customer-reported networking, GPU-interconnect, and node issues on live clusters.
  • Diagnose multi-node NCCL and RDMA performance regressions, including issues sourced to GPUDirect RDMA contention, and separate confirmed findings from working hypotheses when talking to customers.
  • Configure and validate RoCEv2 congestion control (ECN/DCQCN, PFC, DCBX) and reconcile host-side NIC settings against the switch fabric's actual behavior.
  • Diagnose GPU hardware faults using DCGM telemetry and Xid error codes, distinguishing genuine hardware failures (RMA) from driver, firmware, or configuration issues.
  • Own driver and firmware version compatibility across the fleet: GPU driver, NIC firmware (OFED/MOFED), kernel, and validate compatibility before and after updates.
  • Provision, re-provision, and decommission GPU nodes: power-cycling, secure wipe, re-imaging, BIOS/NUMA/PCIe topology validation, and acceptance testing before handoff to a customer.
  • Triage edge-network and jumpbox access issues affecting customer provisioning workflows (e.g., automated blocking of legitimate high-concurrency SSH/Ansible traffic).
  • Write clear incident updates and post-resolution summaries for customer engineering teams, and feed recurring issues back into provisioning defaults and acceptance-test procedures.
Requirements
  • Hands-on experience with RDMA fabrics — RoCEv2 specifically, including ECN/DCQCN, PFC, and DCBX trust-mode configuration on Mellanox/NVIDIA ConnectX NICs.
  • Experience debugging multi-node GPU interconnect performance with NCCL and RDMA benchmarking tools (perftest, nccl-tests).
  • Working knowledge of GPU interconnect topology (NVLink/NVSwitch, NUMA/PCIe affinity) and how it affects distributed training performance.
  • Experience reading GPU health telemetry (NVIDIA DCGM, nvidia-smi, Xid codes) and managing driver/firmware compatibility across a fleet.
  • Strong Linux systems administration, including SSH/host security tooling (fail2ban, sshd) and fleet automation (Ansible or equivalent).
  • Comfortable owning live, customer-facing escalations on production systems, with clear written and verbal communication under time pressure.
Preferred Qualifications
  • Familiarity with native InfiniBand fabrics (OpenSM/UFM) in addition to RoCEv2.
  • Experience with scheduler-level network integration (Slurm, Kubernetes, Run:AI), including SR-IOV and multi-NIC bonding.
  • Prior experience in a vendor-facing escalation role (e.g., NVIDIA/Mellanox support) or as a Field Applications Engineer
What We Offer

Competitive base and strategic ownership in what we believe will be one of the largest infrastructure build-outs of the decade.

Compensation Range

$175,000 - $260,000

Rune is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)

Rune • Mountain View (CA)

Hybrid
USD 175,000 - 260,000
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
GPU Cluster Engineer, Systems & Platform
GPU Cluster Engineer, Systems & Platform

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

Perplexity • New York (NY)

On-site
USD 250,000 - 485,000
Network Engineer, Supercomputing
Network Engineer, Supercomputing

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Unlimited PTO
Health, dental, and vision benefits
Paid parental leave
+1
GPU Cluster Engineer, Networking
GPU Cluster Engineer, Networking

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000
Network Engineer, Supercomputing
Network Engineer, Supercomputing

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited PTO
Paid parental leave
+1
Senior Network Solutions Architect
Senior Network Solutions Architect

Hamilton Barnes ? • San Francisco (CA)

On-site
USD 190,000 - 260,000
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Senior Software Engineer, Cloud-Native Stack – CSP Engagements

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity options
Comprehensive benefits
HPC Operations Engineer
HPC Operations Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Equity eligibility
Diverse work environment
Comprehensive benefits