Senior HPC Support Engineer: Linux, Kubernetes, On-Call

Lambda

United States

On-site

USD 122,000 - 162,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health coverage
Dental
Vision
Wellness stipend
Commuter stipend
401k matched
Flexible PTO

Job summary

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure providing services to researchers, enterprises, and hyperscalers. This role is based in the United States and participates in an on-call rotation to handle the most challenging infrastructure and platform issues.

You will act as a senior escalation point, differentiate among hardware, driver, and kernel problems, and drive permanent fixes by collaborating with engineering and documenting solutions.

Qualifications

  • 3+ years in HPC administration, support, or engineering.
  • Strong Linux system administration experience.
  • Experience with Kubernetes and/or Slurm for cluster orchestration.

Responsibilities

  • Serve as senior escalation point for infra issues down to hardware/kernel level.
  • Differentiate hardware, driver, kernel, and workload misconfig to resolve issues quickly.
  • Identify gaps in processes, tooling, and docs and fix them.
  • Build scripts or small tools using AI-assisted techniques to close gaps.
  • Perform root-cause analysis across distributed systems and GPU clusters.
  • Document solutions and contribute to support procedures.

Skills

HPC administration
Linux admin
Kubernetes
Slurm
CI/CD
AI tooling
Monitoring tools
Kernel debugging
CUDA/NVLink
Networking IB/RoCE
Distributed AI/ML
Cloud networking

Tools

Docker
Terraform
Ansible
Prometheus
Grafana
Datadog
NVIDIA GPU tools

Job description

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure providing services to researchers, enterprises, and hyperscalers. This role is based in the United States and participates in an on-call rotation to handle the most challenging infrastructure and platform issues.

You will act as a senior escalation point, differentiate among hardware, driver, and kernel problems, and drive permanent fixes by collaborating with engineering and documenting solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Support Engineer - GPU Cloud Infra
Senior HPC Support Engineer - GPU Cloud Infra

Neura Market • United States

On-site
USD 150,000 - 190,000
Wellness stipend
Commuter stipend
401k plan with 2% company match (USA)
+1
Senior HPC Support Engineer — Linux, GPUs, Kubernetes
Senior HPC Support Engineer — Linux, GPUs, Kubernetes

Lambda Inc. • United States

On-site
USD 150,000 - 190,000
Health, dental, and vision coverage
401k Plan with 2% company match
Wellness and commuter stipends
+1
HPC Support Engineer
HPC Support Engineer

Lambda • United States

On-site
USD 122,000 - 162,000
Health coverage
Dental
Vision
+4
Senior HPC Validation Engineer – AI Cloud Infrastructure
Senior HPC Validation Engineer – AI Cloud Infrastructure

Lambda • San Jose (CA)

On-site
USD 150,000 - 210,000
Health coverage
Dental coverage
Vision coverage
+4
Senior AI Cloud SRE – HPC & GPU Infra (Hybrid)
Senior AI Cloud SRE – HPC & GPU Infra (Hybrid)

Lambda • United States

Hybrid
USD 140,000 - 200,000
Senior AI Cloud SRE — HPC & GPU Reliability Lead
Senior AI Cloud SRE — HPC & GPU Reliability Lead

Lambda • San Francisco (CA)

On-site
USD 180,000 - 230,000
Health, dental, and vision
401k with company match
Flexible paid time off
+2
Senior AI Cloud SRE — HPC & GPU Infra
Senior AI Cloud SRE — HPC & GPU Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+4
Senior HPC Platform Hardware Engineer - Hybrid (San Jose)
Senior HPC Platform Hardware Engineer - Hybrid (San Jose)

Neura Market • San Jose (CA)

On-site
USD 180,000 - 280,000
Cash & equity compensation
Health, dental, and vision coverage
Wellness stipends
+1
Senior SRE: Managed Kubernetes for AI Cloud Platforms
Senior SRE: Managed Kubernetes for AI Cloud Platforms

Lambda • Bellevue (WA)

On-site
USD 240,000 - 356,000
Health, dental, and vision coverage
401k matching
Flexible PTO
+1
Senior Platform Engineer - AI Cloud Infra (Hybrid)
Senior Platform Engineer - AI Cloud Infra (Hybrid)

Lambda • Bellevue (WA)

Hybrid
USD 190,000 - 270,000
Health, dental, and vision coverage
401k with 2% company match
Flexible PTO
+3