HPC/GPU Systems Engineer

Nscale

Houston, San Francisco, Seattle (TX, CA, WA)

On-site

USD 100,000 - 140,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Nscale is a vertically integrated AI cloud provider delivering end-to-end infrastructure—from energy and data centres to GPU superclusters—across Europe and the US. The Infrastructure Support Engineer role focuses on GPU fleet health, tickets, alarms, and hardware troubleshooting in a fast-paced, customer‑facing environment.

You will own tickets end-to-end, communicate clearly with customers and colleagues, and grow toward Senior through hands-on learning and structured support processes.

Qualifications

  • 3+ years in infrastructure support or support engineering, with customer-facing experience.
  • Solid Linux CLI and troubleshooting skills.
  • Hands-on GPU infrastructure knowledge and hardware troubleshooting experience.
  • Experience with ITSM processes, ticketing, and incident management.
  • Familiarity with GPU/InfiniBand networking concepts is a plus.

Responsibilities

  • Handle day-to-day tickets and alerts in support rotation; escalate when needed.
  • Perform GPU node triage and hardware troubleshooting; interpret nvidia-smi/DCGM outputs and logs.
  • Run diagnostics and hand over clear evidence for vendors; follow runbooks and escalate appropriately.
  • Assist with storage/data-path investigations and gather evidence for senior diagnosis.
  • Document steps, contribute to runbooks, and help improve automation.

Skills

GPU troubleshooting
Linux CLI
Networking basics
ITSM
Scripting
Observability
Communication
Growth mindset
Adaptability
Ticketing

Tools

nvidia-smi
DCGM
BMC
NetBox
mlxlink

Job description

Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.

At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

About the Role

Infrastructure Support Engineers are the delivery engine of Infrastructure Support (L2/L3), handling the day-to-day health of Nscale's GPU fleets — tickets, alerts, hardware faults, and customer issues — across GPU nodes, high-performance networks, Linux, and data centre operations. This is a hands‑on technical role: you'll come in with a strong technical base and working knowledge of GPU infrastructure, and grow toward Senior through exposure to some of the most advanced AI infrastructure in the world.

You will:

  • Own your tickets and tasks end-to-end, escalating early and appropriately when issues exceed your scope — with clean, evidence-rich handovers.
  • Communicate technical detail clearly, specifically, and concisely — in tickets, to customers, and to colleagues. We treat communication quality as a core engineering skill, not a soft skill.
  • Grasp new technical concepts and problems quickly; stay curious and know what questions to ask to get up to speed fast.
  • Bring discipline and organisation: accurate records, structured troubleshooting, reliable follow‑through.
  • Seek feedback and invest in learning — this role is a deliberate pathway to Senior.

Experience required: 3-4+ years in infrastructure support or support engineering roles, including deep support/service desk experience in structured, customer‑facing environments, with working knowledge of GPU infrastructure and hands‑on hardware troubleshooting.

What You’ll Be Doing
  • Join the Support duty rotation and handle day‑to‑day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it.
  • Perform GPU node triage and hardware troubleshooting: interpret nvidia‑smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA.
  • Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and upscale with a handover that lets the next engineer continue without starting from scratch.
  • Assist with storage and data‑path investigations (mounts, connectivity, client‑side symptoms) on high‑performance platforms, gathering evidence for Senior or Engineering‑led diagnosis.
  • Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review.
  • Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels.
  • Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover.
  • Participate in changes under peer review, learning risk assessment and backout practices in live customer environments.
  • Help maintain source‑of‑truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns).
  • Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes.
  • Be the escalation point for onsite DC Operations staff; coordinate smart‑hands tasks within your scope.
  • Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed.
  • Share knowledge by documenting steps you've validated and contributing to training materials. Shadow Seniors during complex work to build capability.
  • Take part in incident reviews as a contributor and help track preventative follow‑ups in your scope.
  • Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed.
  • Participate in on‑call and out‑of‑hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required.
About You
  • Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA‑driven, customer‑facing environments (cloud, data centre, or managed services).
  • Communication. Clear written notes, concise updates, and reliable follow‑through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately.
  • GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands‑on with nvidia‑smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers — reseating components, swap testing, working via BMC/out‑of‑band management — through to preparing RMA evidence. A strong technical base here is required, not a learning goal.
  • Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to elevate.
  • Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high‑performance east‑west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role.
  • Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation.
  • Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review.
  • Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control.
  • Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background.
  • Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior.
  • Adaptability. Able to work in a fast‑moving environment with evolving processes, participate in on‑call after onboarding, and travel when needed.
Nice to Have

These are growth areas, not prerequisites:

  • High‑performance fabrics and GPU‑HPC: exposure to RDMA/InfiniBand, link‑level diagnostics (mlxlink, ibdiagnet, or equivalent), NCCL‑based troubleshooting, or NVLink concepts.
  • High‑performance storage: exposure to VAST or comparable AI‑optimised storage platforms, Ceph, or NFS at scale, including basic storage‑network troubleshooting.
  • OpenStack and fleet operations tooling: familiarity with OpenStack troubleshooting flows, or fleet‑scale tooling for provisioning and health (MAAS, NetBox, Redfish, or similar).
  • Kubernetes: understanding of core concepts (nodes, pods, services, logs) and basic troubleshooting via runbooks. Helpful context for our platform, though not the core of this role.
  • Automation and access tooling: experience with Ansible or Terraform, CI/CD participation (GitHub Actions or similar), or access and security tooling such as Teleport or Vault.
  • Certifications: progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time.
What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.

  • Highly competitive package (base + equity) with reviews every 12 months.
  • Join the fastest‑growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting‑edge AI.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
  • Human‑First Flexibility: We treat you as humans first.
  • Join our thriving remote‑first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work.

We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio‑economic backgrounds.

If there's anything we can do to accommodate your specific situation, please let us know.

The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

The range below reflects the base salary for the position. Actual compensation may vary based on job‑related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programmes. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$100,000 - $140,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:Here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Infrastructure Support Engineer
Senior Infrastructure Support Engineer

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Sr. Staff Security Engineer, Platform Security
Sr. Staff Security Engineer, Platform Security

Socket.dev • Houston (TX)

On-site
USD 210,000 - 250,000
Competitive salary + equity
Bonus + equity programs
Flexible workplace
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
Head of Infrastructure Support
Head of Infrastructure Support

Nscale • San Francisco (CA)

On-site
USD 180,000 - 260,000
Infrastructure Operations Operator (Sauda) EMEA; Norway
Infrastructure Operations Operator (Sauda) EMEA; Norway

Nscale • Town of Norway (WI)

On-site
USD 64,000 - 91,000
Equity
Learning & Development
Flexible work culture
+1
Principal Front-End Network Engineer
Principal Front-End Network Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 260,000
Equity
Competitive salary
Career growth
Principal Front-End Network Engineer
Principal Front-End Network Engineer

Nscale • New York (NY)

On-site
USD 180,000 - 240,000
Equity
Medical, dental, vision insurance
Flexible paid time off
+2
Solutions Engineer
Solutions Engineer

Nscale • Houston (TX)

On-site
USD 140,000 - 193,000
Medical benefits
Dental benefits
Vision benefits
+3
Senior Technical Product Manager, GPU Infrastructure
Senior Technical Product Manager, GPU Infrastructure

Nscale • New York (NY)

On-site
USD 200,000 - 280,000
Medical, dental, vision
Flexible PTO
Retirement plan
Principal Front-End Network Engineer
Principal Front-End Network Engineer

Nscale • Houston (TX)

On-site
USD 150,000 - 260,000
Equity
Medical benefits
Retirement plan
+2