Senior SRE, AI Infrastructure & GPU Fleet Reliability

Hamilton Barnes Associates Limited

San Francisco (CA)

On-site

USD 297,500 - 402,500

Full time

14 days+
Application generator

Get a reply from this recruiter — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Huge stock options
Company bonus
Unlimited PTO
401K + 4% match

Job summary

Hamilton Barnes Associates Limited is seeking a Staff Site Reliability Engineer to lead reliability across large-scale GPU infrastructure used for AI training and inference. You will drive incident response, performance tuning, and platform health, collaborating with cross-functional teams to scale operations.

The role emphasizes hands-on GPU hardware/software expertise, Kubernetes for GPU clusters, and strong programming in Go, Python, or Rust.

Qualifications

  • Hands-on experience operating large-scale GPU infrastructure environments.
  • Staff-level SRE or infrastructure engineering experience supporting mission-critical production systems.
  • Deep expertise with NVIDIA GPU platforms including H100, H200, B200, or GB200 systems.
  • Strong software engineering skills in Go, Python, or Rust.

Responsibilities

  • Lead high-priority incident response across distributed GPU infrastructure environments.
  • Diagnose and resolve issues across the stack including PyTorch, NCCL, CUDA, drivers, networking fabrics, and hardware layers.
  • Own day-to-day operational health of large-scale GPU fleets including lifecycle management and firmware upgrades.
  • Build and maintain observability systems, GPU telemetry platforms, automation tooling, and health-check frameworks.
  • Partner with infrastructure, product, and platform engineering teams to improve reliability and scalability.
  • Mentor engineers across reliability engineering and incident management practices.
  • Contribute to a long-term reliability strategy for hyperscale AI infrastructure.

Skills

GPU infra ops
Staff-level SRE
NVIDIA GPUs
CUDA/NCCL
Go
Python
Rust
Kubernetes
HPC schedulers
Linux systems
Incident response

Tools

Observability tooling

Job description

Hamilton Barnes Associates Limited is seeking a Staff Site Reliability Engineer to lead reliability across large-scale GPU infrastructure used for AI training and inference. You will drive incident response, performance tuning, and platform health, collaborating with cross-functional teams to scale operations.

The role emphasizes hands-on GPU hardware/software expertise, Kubernetes for GPU clusters, and strong programming in Go, Python, or Rust.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE – AI Infra: Scale GPU Clusters & HPC
Senior SRE – AI Infra: Scale GPU Clusters & HPC

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
IPO Equity
10% comapny bonus
401K 4% match
Senior SRE: AI/GPU Scale & Automation
Senior SRE: AI/GPU Scale & Automation

Nscale • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work expectations
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Senior GPU HPC SRE — Remote, Stock Options
Senior GPU HPC SRE — Remote, Stock Options

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Bonus
Remote working option and allowance
Senior SRE: Reliability, Automation & High-Perf Platform
Senior SRE: Reliability, Automation & High-Perf Platform

Hamilton Barnes Associates Limited • New York (NY)

Hybrid
USD 360,000 - 440,000
Strong compensation and bonus potential
Collaborative engineering culture
Work on mission-critical systems
Senior SRE: GPU-Driven, Global Scale & Causal AI
Senior SRE: GPU-Driven, Global Scale & Causal AI

Crossing Hurdles • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior SRE Lead - AI-Driven Reliability & Scale
Senior SRE Lead - AI-Driven Reliability & Scale

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Senior AI Cloud SRE — HPC & GPU Infrastructure
Senior AI Cloud SRE — HPC & GPU Infrastructure

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000