Senior GPU Infrastructure Support Engineer

Nscale

San Francisco (CA)

On-site

USD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Remote-friendly team
Flexible workplace

Job summary

Nscale seeks a Senior Infrastructure Support Engineer to own the health of GPU fleets and high‑performance fabrics. You will operate across GPU hardware, Linux, and data centre operations, bridging Support, DC Operations, and Engineering.

You’ll diagnose complex issues, lead investigations, and implement automation while mentoring mid‑level engineers. Travel to sites may be required, with a remote‑first team structure supporting US and Europe.

Qualifications

  • 6+ years in infrastructure, operations, or support engineering in production environments.
  • 2–3+ years hands‑on with GPU, HPC, or large-scale data centre estates, ideally customer-facing.
  • Strong written discipline; clear incident updates and customer-facing communication.

Responsibilities

  • Join the Support duty rotation as a senior escalation point for incidents and changes.
  • Diagnose GPU node faults across driver, firmware, hardware, and RMA processes.
  • Manage east–west fabrics diagnostics and topology validation on InfiniBand/RoCE fabrics.
  • Investigate data-path issues on high-performance storage platforms and storage-network interactions.
  • Lead structured investigations with root cause analysis and long-term fixes.
  • Author changes in live customer environments with proper risk assessment and backout plans.
  • Improve dashboards, alerts, and runbooks to prevent repeat incidents.

Skills

Infra experience
GPU/HPC knowledge
Linux systems
SRE practices
Incident response
Automation scripting
Network fundamentals
Leadership & mentoring
Customer-facing comms
BMC/Redfish experience

Tools

MAAS
NCCL
Redfish
OpenStack
Prometheus/Grafana
Git
Terraform
Ansible
GitHub Actions

Job description

Nscale seeks a Senior Infrastructure Support Engineer to own the health of GPU fleets and high‑performance fabrics. You will operate across GPU hardware, Linux, and data centre operations, bridging Support, DC Operations, and Engineering.

You’ll diagnose complex issues, lead investigations, and implement automation while mentoring mid‑level engineers. Travel to sites may be required, with a remote‑first team structure supporting US and Europe.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Remote GPU Infrastructure Engineer
Senior Remote GPU Infrastructure Engineer

Nscale • New York (NY)

On-site
USD 120,000 - 170,000
Senior GPU Infra Engineer — Remote
Senior GPU Infra Engineer — Remote

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Remote Infra Operations Engineer - GPU Cloud
Remote Infra Operations Engineer - GPU Cloud

Nscale • Barstow (TX)

Remote
USD 80,000 - 110,000
Competitive pay with equity
Flexible workplace
Dynamic progression plan
Head of GPU Fleet Automation & Infrastructure Engineering
Head of GPU Fleet Automation & Infrastructure Engineering

nscaleoperationsukltd • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale

Luma • Redwood City (CA)

Remote
USD 180,000 - 240,000
Head of Infrastructure Support
Head of Infrastructure Support

Nscale • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Infrastructure Support Engineer
Senior Infrastructure Support Engineer

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)

Rune • Mountain View (CA)

Hybrid
USD 175,000 - 260,000
Senior GPU Compute Infra Architect - Onsite SF
Senior GPU Compute Infra Architect - Onsite SF

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 233,000 - 316,000
Founding-level ownership and visible价值
Direct access to founders
Onsite role in San Francisco
Director, Network & Compute Platform Engineering
Director, Network & Compute Platform Engineering

Nscale • New York (NY)

On-site
USD 240,000 - 320,000
Medical insurance
Dental insurance
Vision care
+3