Senior GPU Infra Engineer — Remote

Nscale

Seattle (WA)

On-site

USD 120,000 - 170,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Remote-first culture
Equity plan
Flexible workplace

Job summary

Nscale is seeking a Senior Infrastructure Support Engineer to own the health of GPU fleets and high-performance fabrics in a hands-on L2/L3 role. You will operate across GPU hardware, Linux, and data centre operations—bridging Support, DC Ops, and Engineering to ensure reliability.

Responsibilities include incident escalation, diagnostics across the full stack, automation, and mentoring mid-level engineers. On-call travel to sites may be required.

Qualifications

  • 6+ years in infrastructure, operations, or support engineering in production environments.
  • 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in customer-facing or escalation-driven capacity.
  • Strong written discipline: notes should let the next engineer pick up where you left off.

Responsibilities

  • Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes.
  • Diagnose and remediate GPU node faults across the full stack—driver, firmware, and hardware layers—from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA.
  • Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics.
  • Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage–network interactions across clients, mounts, VIPs, and routing.
  • Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion.
  • Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans.
  • Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation.
  • Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover.
  • Design and implement automation scripts and small tools to reduce toil and human intervention.
  • Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter.
  • Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews.
  • Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion.
  • Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise.

Skills

GPU HPC expertise
Linux systems engineering
Networking fundamentals
SRE & incident response
Automation scripting
Leadership mentoring
Communication skills
Data centre operations

Tools

nvidia-smi
DCGM
Redfish
mlxlink
ibdiagnet
MAAS

Job description

Nscale is seeking a Senior Infrastructure Support Engineer to own the health of GPU fleets and high-performance fabrics in a hands-on L2/L3 role. You will operate across GPU hardware, Linux, and data centre operations—bridging Support, DC Ops, and Engineering to ensure reliability.

Responsibilities include incident escalation, diagnostics across the full stack, automation, and mentoring mid-level engineers. On-call travel to sites may be required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Infrastructure Support Engineer
Senior GPU Infrastructure Support Engineer

Nscale • San Francisco (CA)

On-site
USD 120,000 - 170,000
Equity
Remote-friendly team
Flexible workplace
GPU Infrastructure Engineer
GPU Infrastructure Engineer

Nscale • Houston (TX), San Francisco (CA), Seattle (WA)

On-site
USD 100,000 - 140,000
Senior GPU Infra Support Engineer (Remote, Tier 2/3)
Senior GPU Infra Support Engineer (Remote, Tier 2/3)

Hydra Host, Inc. • Miami (FL), Northern (KY)

Hybrid
USD 90,000 - 130,000
Senior InfraOps Engineer — GPU Infra, Hybrid, Equity
Senior InfraOps Engineer — GPU Infra, Hybrid, Equity

Lightning AI • New York (NY)

Hybrid
USD 160,000 - 200,000
Health coverage
Equity/RSUs
401(k) matching
+1
Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Technical Lead - GPU Infrastructure (Fully Remote)
Technical Lead - GPU Infrastructure (Fully Remote)

Jobgether SRL • United States

Remote
USD 170,000 - 250,000
GPU Infrastructure Support Engineer - Tier 2/3 (Remote)
GPU Infrastructure Support Engineer - Tier 2/3 (Remote)

Hydra Host • Miami (FL)

Remote
USD 90,000 - 130,000
Senior GPU Fleet Operations Product Manager
Senior GPU Fleet Operations Product Manager

Nscale • New York (NY)

On-site
USD 200,000 - 280,000
Medical, dental, vision
Flexible PTO
Retirement plan
Senior GPU Compute Fleet Engineer | Remote-First
Senior GPU Compute Fleet Engineer | Remote-First

Boundless Networks • United States

Remote
USD 175,000 - 250,000
Equity
Health, dental, vision
Flexible PTO
+2
Senior GPU Infra Lead: Slurm, Kubernetes & Platform
Senior GPU Infra Lead: Slurm, Kubernetes & Platform

Jobgether SRL • United States

Remote
USD 170,000 - 250,000