Senior HPC Support Engineer — Linux, GPUs, Kubernetes

Lambda Inc.

United States

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision coverage
401k Plan with 2% company match
Wellness and commuter stipends
Flexible paid time off

Job summary

Lambda is seeking a senior HPC/Linux engineer to serve as a primary escalation point for infrastructure issues across GPU clusters. You will distinguish root causes, implement fixes, and drive improvements with engineering teams.

You’ll mentor junior staff and participate in rotating on-call duties, with emphasis on robust, scalable Linux systems and CUDA-enabled workloads. The role requires extensive experience with Kubernetes/Slurm, CI/CD, and monitoring tools, plus strong scripting and

Qualifications

  • This role requires 3+ years in HPC/Linux environments with hands-on administration.
  • Experience with Kubernetes or Slurm for cluster orchestration is preferred.
  • Strong coding ability, CI/CD experience, and use of AI-assisted tooling to move fast.
  • Proficient in monitoring/logging (Prometheus, Grafana, Datadog) and kernel-level debugging.
  • Experience with CUDA, NCCL, NVLink, GPUDirect RDMA and high-throughput networks.
  • Knowledge of distributed AI/HPC workloads and cloud networking (TCP/IP, VPN, firewalls).
  • Ability to mentor junior engineers and participate in on-call rotations.

Responsibilities

  • Serve as senior escalation point for infrastructure issues down to hardware/drivers.
  • Differentiate hardware vs driver vs kernel vs workload misconfigurations clearly.
  • Identify gaps in processes/tools/docs and implement fixes.
  • Develop small internal tools and scripts using AI-assisted workflows.
  • Conduct root-cause analysis across distributed GPU clusters.
  • Document solutions and improve support procedures.
  • Collaborate with engineering to convert pain points into fixes.
  • Mentor junior engineers and take escalations during on-call rotation.
  • Lead major incident ownership during on-call, with quick resolution.
  • Pitch in during high-volume deployments when needed.

Skills

HPC administration
Linux system administration
Kernel debugging
CI/CD pipelines
Scripting with AI tools
Monitoring & logging
CUDA/NCCL/GPUDirect
NVLink & InfiniBand
TCP/IP/VPN/Firewall
On-call support & mentoring
Distributed AI/HPC workloads
Kubernetes/Grid orchestration

Tools

Kubernetes
Slurm
Docker
Terraform
Ansible

Job description

Lambda is seeking a senior HPC/Linux engineer to serve as a primary escalation point for infrastructure issues across GPU clusters. You will distinguish root causes, implement fixes, and drive improvements with engineering teams.

You’ll mentor junior staff and participate in rotating on-call duties, with emphasis on robust, scalable Linux systems and CUDA-enabled workloads. The role requires extensive experience with Kubernetes/Slurm, CI/CD, and monitoring tools, plus strong scripting and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Support Engineer - GPU/Kernel Master
Senior HPC Support Engineer - GPU/Kernel Master

Lambda Labs • United States

On-site
USD 140,000 - 200,000
Health, dental, vision coverage
401k plan with 2% company match
Wellness & commuter stipends
Senior HPC Support Engineer - GPU Cloud Infra
Senior HPC Support Engineer - GPU Cloud Infra

Neura Market • United States

On-site
USD 150,000 - 190,000
Wellness stipend
Commuter stipend
401k plan with 2% company match (USA)
+1
Senior HPC Support Engineer: Linux, Kubernetes, On-Call
Senior HPC Support Engineer: Linux, Kubernetes, On-Call

Lambda • United States

On-site
USD 122,000 - 162,000
Health coverage
Dental
Vision
+4
HPC Support Engineer
HPC Support Engineer

Lambda Labs • United States

On-site
USD 140,000 - 200,000
Health, dental, vision coverage
401k plan with 2% company match
Wellness & commuter stipends
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters

Parallel Works • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
HPC Support Engineer
HPC Support Engineer

Lambda • United States

On-site
USD 122,000 - 162,000
Health coverage
Dental
Vision
+4
Senior Cloud Platform Engineer, GPU Core & Lifecycle
Senior Cloud Platform Engineer, GPU Core & Lifecycle

Lambda Labs • United States

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+5
Senior HPC Engineer - Linux, Kubernetes & GPUs
Senior HPC Engineer - Linux, Kubernetes & GPUs

SpaceX • Brownsville (TX)

On-site
USD 140,000 - 200,000
Senior Cloud Infrastructure Engineer – GPU & DPU
Senior Cloud Infrastructure Engineer – GPU & DPU

Lambda • United States

Remote
USD 180,000 - 260,000
Senior HPC Systems Administrator - Linux, Slurm, GPU
Senior HPC Systems Administrator - Linux, Slurm, GPU

Jahnel Group • United States

On-site
USD 120,000 - 180,000