Senior HPC Support Engineer - GPU Cloud Infra

Neura Market

United States

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Wellness stipend
Commuter stipend
401k plan with 2% company match (USA)
Flexible paid time off

Job summary

Lambda, The Superintelligence Cloud, is seeking a senior HPC-focused engineer to join our on-call rotation and directly support our Linux-based GPU infrastructure. You will serve as a high‑level escalation point, diagnose issues down to hardware, driver, or kernel, and work with engineering to implement permanent fixes.

You have 3+ years of hands-on HPC experience, strong Linux administration skills, and familiarity with Kubernetes or Slurm for cluster orchestration, plus CUDA/NCCL/NVLink.

Qualifications

  • 3+ years of hands-on HPC experience in an administration, support, or engineering role.
  • Very strong understanding and experience supporting Linux in a system administration role.
  • Proven experience in HPC environments, showcasing Linux cluster administration with Kubernetes/Slurm.
  • Strong coding ability and CI/CD experience, with AI-assisted tooling experience.
  • Proficiency with monitoring/logging tools (Prometheus, Grafana, Datadog).
  • Strong log analysis, kernel debugging, and performance profiling skills.
  • Experience with CUDA, NCCL, NVLink, GPUDirect RDMA.
  • Experience with high throughput networking technologies (IB/RoCE).
  • Knowledge of distributed AI/ML or HPC workloads.
  • Knowledge of TCP/IP, VPN, and firewalls in cloud environments.
  • Ability to work independently and mentor junior engineers.

Responsibilities

  • Serve as a senior escalation point, troubleshooting the hardest infra issues down to hardware or kernel.
  • Distinguish hardware, driver, kernel, and workload misconfigurations to resolve issues quickly.
  • Identify gaps in processes, tooling, and docs and implement fixes.
  • Use AI tools to build scripts and small internal tools to close gaps.
  • Perform root-cause analysis across distributed systems and GPU infra.
  • Document solutions and evolve support procedures.
  • Collaborate with engineering to turn recurring pain points into permanent fixes.
  • Mentor peers and take escalations from teammates.
  • Participate in rotating on-call and own major incidents.
  • Pitch in during fast, high-volume deployments as needed.

Skills

HPC administration
Linux system administration
Kubernetes
Slurm
CI/CD
Monitoring (Prometheus)
Logging (Datadog)
CUDA/NCCL/NVLink
GPUDirect RDMA
Networking (IB/RoCE)
Distributed AI/ML workloads
TCP/IP/VPN/firewalls
Mentoring

Education

Bachelor's degree in CS/Engineering

Tools

Docker
Terraform
Ansible

Job description

Lambda, The Superintelligence Cloud, is seeking a senior HPC-focused engineer to join our on-call rotation and directly support our Linux-based GPU infrastructure. You will serve as a high‑level escalation point, diagnose issues down to hardware, driver, or kernel, and work with engineering to implement permanent fixes.

You have 3+ years of hands-on HPC experience, strong Linux administration skills, and familiarity with Kubernetes or Slurm for cluster orchestration, plus CUDA/NCCL/NVLink.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Support Engineer — Linux, GPUs, Kubernetes
Senior HPC Support Engineer — Linux, GPUs, Kubernetes

Lambda Inc. • United States

On-site
USD 150,000 - 190,000
Health, dental, and vision coverage
401k Plan with 2% company match
Wellness and commuter stipends
+1
Senior HPC Support Engineer - GPU/Kernel Master
Senior HPC Support Engineer - GPU/Kernel Master

Lambda Labs • United States

On-site
USD 140,000 - 200,000
Health, dental, vision coverage
401k plan with 2% company match
Wellness & commuter stipends
Senior HPC Support Engineer: Linux, Kubernetes, On-Call
Senior HPC Support Engineer: Linux, Kubernetes, On-Call

Lambda • United States

On-site
USD 122,000 - 162,000
Health coverage
Dental
Vision
+4
Senior Cloud Infrastructure Engineer – GPU & DPU
Senior Cloud Infrastructure Engineer – GPU & DPU

Lambda • United States

Remote
USD 180,000 - 260,000
Senior Cloud Platform Engineer - GPU Infra, Hybrid
Senior Cloud Platform Engineer - GPU Infra, Hybrid

Neura Market • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Wellness stipend
Commuter stipend
401k with company match
Senior Cloud Platform Engineer - GPU Infrastructure
Senior Cloud Platform Engineer - GPU Infrastructure

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1
HPC Support Engineer
HPC Support Engineer

Lambda • United States

On-site
USD 122,000 - 162,000
Health coverage
Dental
Vision
+4
HPC Support Engineer
HPC Support Engineer

Lambda Labs • United States

On-site
USD 140,000 - 200,000
Health, dental, vision coverage
401k plan with 2% company match
Wellness & commuter stipends
Senior Cloud Platform Engineer, GPU Core & Lifecycle
Senior Cloud Platform Engineer, GPU Core & Lifecycle

Lambda Labs • United States

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+5
Senior HPC Systems Architect: Liquid-Cooled GPU Cloud
Senior HPC Systems Architect: Liquid-Cooled GPU Cloud

Socket.dev • San Jose (CA)

Hybrid
USD 180,000 - 260,000