Senior GPU Infra Engineer — Remote

Nscale

Seattle (WA)

On-site

USD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote-first culture
Equity plan
Flexible workplace

Job summary

Nscale is seeking a Senior Infrastructure Support Engineer to own the health of GPU fleets and high-performance fabrics in a hands-on L2/L3 role. You will operate across GPU hardware, Linux, and data centre operations—bridging Support, DC Ops, and Engineering to ensure reliability.

Responsibilities include incident escalation, diagnostics across the full stack, automation, and mentoring mid-level engineers. On-call travel to sites may be required.

Qualifications

  • 6+ years in infrastructure, operations, or support engineering in production environments.
  • 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in customer-facing or escalation-driven capacity.
  • Strong written discipline: notes should let the next engineer pick up where you left off.

Responsibilities

  • Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes.
  • Diagnose and remediate GPU node faults across the full stack—driver, firmware, and hardware layers—from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA.
  • Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics.
  • Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage–network interactions across clients, mounts, VIPs, and routing.
  • Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion.
  • Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans.
  • Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation.
  • Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover.
  • Design and implement automation scripts and small tools to reduce toil and human intervention.
  • Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter.
  • Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews.
  • Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion.
  • Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise.

Skills

GPU HPC expertise
Linux systems engineering
Networking fundamentals
SRE & incident response
Automation scripting
Leadership mentoring
Communication skills
Data centre operations

Tools

nvidia-smi
DCGM
Redfish
mlxlink
ibdiagnet
MAAS

Job description

Nscale is seeking a Senior Infrastructure Support Engineer to own the health of GPU fleets and high-performance fabrics in a hands-on L2/L3 role. You will operate across GPU hardware, Linux, and data centre operations—bridging Support, DC Ops, and Engineering to ensure reliability.

Responsibilities include incident escalation, diagnostics across the full stack, automation, and mentoring mid-level engineers. On-call travel to sites may be required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Remote GPU Infrastructure Engineer
Senior Remote GPU Infrastructure Engineer

Nscale • New York (NY)

On-site
USD 120,000 - 170,000
Senior GPU Infrastructure Support Engineer
Senior GPU Infrastructure Support Engineer

Nscale • San Francisco (CA)

On-site
USD 120,000 - 170,000
Equity
Remote-friendly team
Flexible workplace
Remote Infra Operations Engineer - GPU Cloud
Remote Infra Operations Engineer - GPU Cloud

Nscale • Barstow (TX)

Remote
USD 80,000 - 110,000
Competitive pay with equity
Flexible workplace
Dynamic progression plan
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale

Luma • Redwood City (CA)

Remote
USD 180,000 - 240,000
Senior GPU Infra Reliability Engineer - Remote
Senior GPU Infra Reliability Engineer - Remote

Luma AI • United States

Remote
USD 180,000 - 240,000
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)

Rune • Mountain View (CA)

Hybrid
USD 175,000 - 260,000
Head of GPU Fleet Automation & Infrastructure Engineering
Head of GPU Fleet Automation & Infrastructure Engineering

nscaleoperationsukltd • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior GPU Compute Infra Architect - Onsite SF
Senior GPU Compute Infra Architect - Onsite SF

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 233,000 - 316,000
Founding-level ownership and visible价值
Direct access to founders
Onsite role in San Francisco
Senior GPU Compute Fleet Engineer | Remote-First
Senior GPU Compute Fleet Engineer | Remote-First

Boundless Networks • United States

Remote
USD 175,000 - 250,000
Equity
Health, dental, vision
Flexible PTO
+2