Senior HPC Support Engineer - GPU/Kernel Master

Lambda Labs

United States

On-site

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, vision coverage
401k plan with 2% company match
Wellness & commuter stipends

Job summary

Lambda Labs is seeking a Senior HPC Systems Engineer to join our on-site team in the United States. You will act as a senior escalation point, troubleshooting complex GPU infrastructure issues down to hardware and kernel levels, while mentoring junior engineers.

We value strong Linux administration, HPC orchestration with Kubernetes/Slurm, and CI/CD practices. You will craft documentation, identify gaps, and contribute to fast deployments with AI-assisted tooling.

Qualifications

  • 3+ years of hands-on HPC experience in admin, support, or engineering.
  • Strong Linux system administration experience.
  • Experience supporting HPC environments with Kubernetes and/or Slurm.
  • Solid CI/CD background and use of AI-assisted tools.
  • Proficiency with monitoring/logging tools: Prometheus, Grafana, Datadog.
  • Kernel debugging and performance profiling skills.
  • CUDA/NVIDIA GPUs and Infiniband knowledge.
  • TCP/IP, VPN, and cloud firewall familiarity.
  • Ability to mentor junior engineers.

Responsibilities

  • Serve as a senior escalation point for infrastructure and platform issues to hardware/kernel levels.
  • Distinguish hardware, driver, kernel, and workload configuration problems to resolve correctly.
  • Identify and fix tooling, process, and documentation gaps proactively.
  • Develop scripts or small tools with AI assistance to close operational gaps.
  • Perform root-cause analysis across distributed GPU infrastructure.
  • Document solutions and evolve support procedures.
  • Collaborate with engineering to convert recurring pain points into permanent fixes.
  • Mentor and train junior support engineers.
  • Participate in rotating on-call and own major incidents.
  • Contribute during fast, high-volume deployments.

Skills

HPC experience
Linux administration
Kubernetes
Slurm
CI/CD
AI-assisted tooling
Monitoring tools
Kernel debugging
Networking
CUDA/NVIDIA GPUs
Infiniband

Tools

Docker
Kubernetes
Terraform
Ansible
Prometheus
Grafana
Datadog

Job description

Lambda Labs is seeking a Senior HPC Systems Engineer to join our on-site team in the United States. You will act as a senior escalation point, troubleshooting complex GPU infrastructure issues down to hardware and kernel levels, while mentoring junior engineers.

We value strong Linux administration, HPC orchestration with Kubernetes/Slurm, and CI/CD practices. You will craft documentation, identify gaps, and contribute to fast deployments with AI-assisted tooling.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Support Engineer — Linux, GPUs, Kubernetes
Senior HPC Support Engineer — Linux, GPUs, Kubernetes

Lambda Inc. • United States

On-site
USD 150,000 - 190,000
Health, dental, and vision coverage
401k Plan with 2% company match
Wellness and commuter stipends
+1
Senior HPC Support Engineer - GPU Cloud Infra
Senior HPC Support Engineer - GPU Cloud Infra

Neura Market • United States

On-site
USD 150,000 - 190,000
Wellness stipend
Commuter stipend
401k plan with 2% company match (USA)
+1
Senior HPC Support Engineer: Linux, Kubernetes, On-Call
Senior HPC Support Engineer: Linux, Kubernetes, On-Call

Lambda • United States

On-site
USD 122,000 - 162,000
Health coverage
Dental
Vision
+4
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters
Senior HPC Systems Engineer — Secure Hybrid GPU Clusters

Parallel Works • Chicago (IL)

Hybrid
USD 140,000 - 190,000
Medical, vision, dental coverage
401(k) with company match
Short term disability
+1
Senior Cloud Infrastructure Engineer – GPU & DPU
Senior Cloud Infrastructure Engineer – GPU & DPU

Lambda • United States

Remote
USD 180,000 - 260,000
Senior Cloud Platform Engineer, GPU Core & Lifecycle
Senior Cloud Platform Engineer, GPU Core & Lifecycle

Lambda Labs • United States

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+5
Senior HPC Hardware Platform Lead for AI Cloud
Senior HPC Hardware Platform Lead for AI Cloud

Lambda Labs • United States

On-site
USD 150,000 - 230,000
Health insurance
Dental and vision coverage
Wellness stipend
+3
Senior HPC Systems Administrator - Linux, Slurm, GPU
Senior HPC Systems Administrator - Linux, Slurm, GPU

Jahnel Group • United States

On-site
USD 120,000 - 180,000
Senior HPC Systems Architect – Liquid-Cooled AI GPU Infra
Senior HPC Systems Architect – Liquid-Cooled AI GPU Infra

Lambda Labs • United States

Hybrid
USD 180,000 - 280,000
Equity compensation
Health, dental and vision coverage
401k with 2% company match
+1
HPC Support Engineer
HPC Support Engineer

Lambda Labs • United States

On-site
USD 140,000 - 200,000
Health, dental, vision coverage
401k plan with 2% company match
Wellness & commuter stipends