Senior Infrastructure Support Engineer

Nscale

Seattle (WA)

On-site

USD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote-first culture
Equity plan
Flexible workplace

Job summary

Nscale is seeking a Senior Infrastructure Support Engineer to own the health of GPU fleets and high-performance fabrics in a hands-on L2/L3 role. You will operate across GPU hardware, Linux, and data centre operations—bridging Support, DC Ops, and Engineering to ensure reliability.

Responsibilities include incident escalation, diagnostics across the full stack, automation, and mentoring mid-level engineers. On-call travel to sites may be required.

Qualifications

  • 6+ years in infrastructure, operations, or support engineering in production environments.
  • 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in customer-facing or escalation-driven capacity.
  • Strong written discipline: notes should let the next engineer pick up where you left off.

Responsibilities

  • Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes.
  • Diagnose and remediate GPU node faults across the full stack—driver, firmware, and hardware layers—from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA.
  • Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics.
  • Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage–network interactions across clients, mounts, VIPs, and routing.
  • Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion.
  • Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans.
  • Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation.
  • Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover.
  • Design and implement automation scripts and small tools to reduce toil and human intervention.
  • Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter.
  • Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews.
  • Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion.
  • Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise.

Skills

GPU HPC expertise
Linux systems engineering
Networking fundamentals
SRE & incident response
Automation scripting
Leadership mentoring
Communication skills
Data centre operations

Tools

nvidia-smi
DCGM
Redfish
mlxlink
ibdiagnet
MAAS

Job description

About Nscale

Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack—energy, data centres, GPU superclusters, orchestration, and AI services—delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.

About The Role (Job Purpose)

Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands‑on L2/L3 role operating at the intersection of GPU hardware, east‑west networking, Linux, and data centre operations—acting as the operational bridge between Support, DC Operations, and Engineering.

You Will
  • Own complex, ambiguous problems end-to-end and make decisive calls in a results‑driven environment, taking calculated risks where speed matters.
  • Communicate technical detail clearly, specifically, and concisely— to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill.
  • Influence without authority and build strong relationships with senior stakeholders across the business to get things done.
  • Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast.
  • Bring discipline and organisation: evidence‑led investigations, accurate records, clean handovers.
What You’ll Be Doing (Responsibilities)
  • Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes.
  • Diagnose and remediate GPU node faults across the full stack—driver, firmware, and hardware layers—from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA.
  • Own east‑west fabric health: run link-level diagnostics (mlxlink, ibdiagnet or equivalent), isolate transceiver, optics, cabling, and switch‑port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics.
  • Investigate data-path issues on high‑performance storage platforms (e.g. VAST), including storage–network interactions across clients, mounts, VIPs, and routing.
  • Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion.
  • Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans.
  • Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation.
  • Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover.
  • Design and implement automation scripts and small tools to reduce toil and human intervention.
  • Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter.
  • Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews.
  • Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion.
  • Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise.
About You (Skills / Qualifications Experience)
  • Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2–3+ years hands‑on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity.
  • Communication. Able to explain complex technical detail clearly, specifically, and concisely— in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch.
  • GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA.
  • High-performance east-west fabrics. Hands‑on experience with RDMA fabrics—InfiniBand and/or RoCE—including link-layer diagnostics (mlxlink, ibdiagnet or equivalent), transceiver and cabling fault isolation, and understanding of rail‑optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters.
  • HPC scheduling. Slurm operations for large multi-GPU jobs—containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures.
  • Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production.
  • Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets.
  • Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south.
  • Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews.
  • Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans.
  • SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools.
  • Automation and Git. Scripting skills in Bash, Python or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform or similar).
  • Data centre fundamentals. Understanding of how data centres operate—servers, networks, storage, power, and cooling—ideally gained through an operational support background.
  • Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve.
  • Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work.
Nice to Have
  • High-performance storage. Hands‑on experience with VAST or comparable AI-optimised storage platforms, or Ceph/parallel filesystems and NFS at scale (multipath, remoteports, nconnect), including diagnosing storage–network interaction and data-path performance issues.
  • OpenStack and fleet operations tooling. OpenStack operations experience (Neutron, Cinder, error triage), plus familiarity with fleet-scale tooling for provisioning, health, and remediation across large GPU estates (MAAS, NetBox, Redfish-driven automation or similar).
  • Kubernetes. Operating and troubleshooting clusters, including GPU operator stacks and understanding how physical resources are abstracted up the stack. Helpful context for our platform, though not the core of this role.
  • Automation at scale. Automated network configuration with safe, repeatable changes in business-critical environments; GitOps and CI/CD pipelines (GitHub Actions or similar); access and security tooling such as Teleport or Vault in production.
  • Certifications. Relevant GPU/HPC, datacenter architecture, Linux, networking, Kubernetes, cloud, or security certifications (e.g. RHCSA/RHCE, CKA, NVIDIA-certified) are a plus.
What We Can Offer You
  • Highly competitive package, including base salary and equity, with reviews every 12 months.
  • Join one of the fastest-growing tech startups: your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
  • Human-first flexibility. We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
  • Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you wherever you work.
Equal Opportunities Statement

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there’s anything we can do to accommodate your specific situation, please let us know.

Salary Range: $120,000 USD – $170,000 USD

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Operations Engineer Greensboro, NC
Infrastructure Operations Engineer Greensboro, NC

Nscale • Winston-Salem (NC)

Hybrid
USD 100,000 - 160,000
Competitive package (base + equity)
Flexible workplace and support for personal growth
Medical, dental, and vision benefits
Support Desk Engineer
Support Desk Engineer

Nscale • United States

Remote
USD 50,000 - 100,000
Competitive package with reviews every 12 months
Dynamic progression plan tailored to ambitions
Human-first flexibility in a remote-first environment
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Equity
Base salary + equity
Career progression
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • Houston (TX)

On-site
USD 150,000 - 215,000
Base salary + equity
Annual reviews
Growth opportunities
Senior Software Engineering Manager - Fleet Management
Senior Software Engineering Manager - Fleet Management

Socket.dev • Seattle (WA)

On-site
USD 300,000 - 350,000
Equity
Bonus potential
Flexible work policy
+1
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • New York (NY)

On-site
USD 150,000 - 215,000
Equity
Medical insurance
Dental insurance
+4
Infrastructure Software Engineer, Fleet & Automation Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation Houston; New York; San Francisco; Seattle

Nscale • Northern (KY)

Hybrid
USD 150,000 - 215,000
Base + equity
Equity
Growth opportunities
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Equity
Competitive compensation
Benefits package (medical, dental, V)
Senior Software Engineering Manager - Fleet Management
Senior Software Engineering Manager - Fleet Management

Nscale • Bellevue (WA)

On-site
USD 300,000 - 350,000
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities