Senior Remote GPU Infrastructure Engineer

Nscale

New York (NY)

On-site

USD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nscale seeks a Senior Infrastructure Support Engineer to own GPU fleet health and high‑performance fabrics. This hands‑on L2/L3 role bridges GPU hardware, Linux, networking, and data centre operations, connecting Support, DC Operations, and Engineering.

You will drive incidents end‑to‑end, perform root‑cause analyses, implement automation, and mentor engineers. Travel to Nscale or customer sites may be required for onsite expertise.

Qualifications

  • 6+ years in infrastructure, operations, or support engineering in production environments.
  • 2–3+ years hands‑on with GPU, HPC, or large‑scale data centre estates.
  • Strong written and spoken communication across customers and stakeholders.
  • Deep understanding of GPU drivers, firmware, and runtime stacks for AI clusters.
  • Experience with east‑west fabric diagnostics and high‑speed networks.
  • Proven ability to author/run changes with risk assessment and backout plans.

Responsibilities

  • Join and lead the Support duty rotation as a senior escalation point.
  • Diagnose GPU node faults across driver, firmware, and hardware layers.
  • Own east‑west fabric health; diagnose and isolate faults in InfiniBand/RoCE fabrics.
  • Investigate data-path issues on high‑performance storage platforms.
  • Perform structured, hypothesis‑driven investigations and RCA.
  • Author and execute changes in live customer environments with proper risk assessment.
  • Improve dashboards, alerts, and runbooks; drive automation to reduce toil.
  • Record and communicate status with clear customer‑impact notes.
  • Design automation scripts and small tools to reduce manual effort.
  • Be a primary escalation point; mentor mid‑level engineers and share knowledge.
  • Lead with disciplined decision making; travel onsite as needed.

Skills

GPU & HPC expertise
Communication skills
GPU drivers & DCGM
East–west fabric debugging
Linux systems engineering
Networking fundamentals
Observability & incident response
Automation scripting
Change & risk assessment
Mentoring engineers
Leadership under pressure
Travel for onsite work

Tools

nvidia-smi / DCGM tooling
mlxlink / ibdiagnet
Redfish / BMC
MAAS
Pyxis / Enroot
Slurm
Prometheus / Grafana
Terraform / Ansible
Git / GitOps
OpenStack tooling

Job description

Nscale seeks a Senior Infrastructure Support Engineer to own GPU fleet health and high‑performance fabrics. This hands‑on L2/L3 role bridges GPU hardware, Linux, networking, and data centre operations, connecting Support, DC Operations, and Engineering.

You will drive incidents end‑to‑end, perform root‑cause analyses, implement automation, and mentor engineers. Travel to Nscale or customer sites may be required for onsite expertise.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infra Engineer — Remote
Senior GPU Infra Engineer — Remote

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Senior GPU Infrastructure Support Engineer
Senior GPU Infrastructure Support Engineer

Nscale • San Francisco (CA)

On-site
USD 120,000 - 170,000
Equity
Remote-friendly team
Flexible workplace
Remote Infra Operations Engineer - GPU Cloud
Remote Infra Operations Engineer - GPU Cloud

Nscale • Barstow (TX)

Remote
USD 80,000 - 110,000
Competitive pay with equity
Flexible workplace
Dynamic progression plan
Director, Network & Compute Platform Engineering
Director, Network & Compute Platform Engineering

Nscale • San Francisco (CA)

On-site
USD 240,000 - 320,000
Bonus eligibility
Equity
Retirement plan
Director, Network & Compute Platform Engineering
Director, Network & Compute Platform Engineering

Nscale • New York (NY)

On-site
USD 240,000 - 320,000
Medical insurance
Dental insurance
Vision care
+3
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale

Luma • Redwood City (CA)

Remote
USD 180,000 - 240,000
Head of GPU Fleet Automation & Infrastructure Engineering
Head of GPU Fleet Automation & Infrastructure Engineering

nscaleoperationsukltd • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)

Rune • Mountain View (CA)

Hybrid
USD 175,000 - 260,000
Senior Network Engineer, AI Infra & High-Performance Cloud
Senior Network Engineer, AI Infra & High-Performance Cloud

Nscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Medical insurance
Retirement plan
Flexible PTO
Director of Network & Compute Platforms
Director of Network & Compute Platforms

Nscale • Houston (TX)

On-site
USD 240,000 - 320,000