Senior AI Cloud SRE — HPC & GPU Infrastructure

Neura Market

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lambda is seeking an experienced Site Reliability Engineer to monitor and optimize AI compute clusters. You will deploy, scale, and automate HPC environments, building reliable operations for GPU-rich workloads across on-site teams.

You will implement robust tooling in Python/Go, use Prometheus and Grafana for observability, and drive incident response with a focus on stability and efficiency. This role requires presence in SF or Bellevue with a hybrid setup.

Qualifications

  • 7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar role.
  • Strong understanding of modern AI infrastructure, from GPU architectures to hardware performance optimization.
  • Strong Linux skills in distributed environments.
  • Experience configuring and troubleshooting InfiniBand (IB), RoCE, CLOS fabrics, 100GbE, Ethernet switching, GPU-direct, and NCCL.
  • Proficiency in Python and Go with internal tooling improvements.
  • Experience with monitoring and alerting tools (Prometheus, Grafana, Clickhouse).
  • Proficiency with automation/configuration tools (Ansible, Terraform).
  • Excellent problem-solving and attention to detail; continuous improvement mindset.

Responsibilities

  • Build and operate monitoring and alerting for cluster health across fabric, GPU, power/thermal, and job signals.
  • Remotely deploy and configure large-scale HPC clusters for AI workloads with automation.
  • Automate cluster lifecycle: OS, firmware, drivers, networking as code (Ansible, Terraform).
  • Create runbooks and automated remediations for common cluster failures for safe execution by Support teams.
  • Troubleshoot cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric, and switches with on-site teams.
  • Lead incident responses during on-call rotations.
  • Contribute to SOPs and provide clear requirements to engineering for stability and efficiency.

Skills

SRE
HPC Engineering
DevOps
Python
Go
Monitoring

Tools

Prometheus
Grafana
Clickhouse

Job description

Lambda is seeking an experienced Site Reliability Engineer to monitor and optimize AI compute clusters. You will deploy, scale, and automate HPC environments, building reliable operations for GPU-rich workloads across on-site teams.

You will implement robust tooling in Python/Go, use Prometheus and Grafana for observability, and drive incident response with a focus on stability and efficiency. This role requires presence in SF or Bellevue with a hybrid setup.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Cloud SRE — HPC & GPU Infra
Senior AI Cloud SRE — HPC & GPU Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+4
Senior AI Cloud SRE — HPC & GPU Reliability Lead
Senior AI Cloud SRE — HPC & GPU Reliability Lead

Lambda • San Francisco (CA)

On-site
USD 180,000 - 230,000
Health, dental, and vision
401k with company match
Flexible paid time off
+2
Senior AI Cloud SRE — Hybrid (SF/Bellevue)
Senior AI Cloud SRE — Hybrid (SF/Bellevue)

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Health, dental, vision coverage
Wellness stipend
Commuter stipend
+2
Senior AI Cloud SRE – HPC & GPU Infra (Hybrid)
Senior AI Cloud SRE – HPC & GPU Infra (Hybrid)

Lambda • United States

Hybrid
USD 140,000 - 200,000
Senior Kubernetes SRE — Scale AI Clusters & Automation
Senior Kubernetes SRE — Scale AI Clusters & Automation

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity compensation
Health, dental and vision coverage
Wellness and commuter stipends
+2
Senior SRE: Managed Kubernetes for AI Cloud
Senior SRE: Managed Kubernetes for AI Cloud

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Site Reliability Engineer - Fleet
Senior Site Reliability Engineer - Fleet

Lambda • United States

Hybrid
USD 140,000 - 200,000
Senior SRE: Managed Kubernetes for AI Cloud (Hybrid)
Senior SRE: Managed Kubernetes for AI Cloud (Hybrid)

Lambda • San Francisco (CA)

Hybrid
USD 170,000 - 260,000
401k Plan with company match (USA)
Health, dental, and vision coverage
Wellness and commuter stipends
+1
Senior SRE: Managed Kubernetes for AI Cloud Platforms
Senior SRE: Managed Kubernetes for AI Cloud Platforms

Lambda • Bellevue (WA)

On-site
USD 240,000 - 356,000
Health, dental, and vision coverage
401k matching
Flexible PTO
+1
Senior SRE: AI Cloud Platform & Kubernetes Expert
Senior SRE: AI Cloud Platform & Kubernetes Expert

Lambda • Bellevue (WA)

On-site
USD 180,000 - 260,000
Health insurance
Dental insurance
Vision insurance
+3