Senior SRE – AI Infra: Scale GPU Clusters & HPC

Hamilton Barnes Associates Limited

San Francisco (CA)

On-site

USD 225,000 - 275,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

IPO Equity
10% comapny bonus
401K 4% match

Job summary

Hamilton Barnes Associates Limited is seeking an experienced SRE/Infrastructure Engineer to help build a seed-stage AI infrastructure company with large-scale GPU clusters for training and inference. You design, deploy, and maintain compute environments and ensure reliability across Slurm and Kubernetes.

You will implement automation, observability, and IaC practices, collaborating with ML and platform teams to optimize GPU utilization, data flow, and latency.

Qualifications

  • 7+ years in SRE/DevOps or infrastructure engineering for large-scale compute.
  • Hands-on with Kubernetes and Slurm for cluster orchestration.
  • Strong Linux, networking, and NVIDIA GPU infra knowledge.
  • Proficient in Python, Go, or Bash for automation.
  • Experience with observability stacks and incident response.
  • Familiar with HPC/AI training infra at scale.
  • Background in reliability engineering or distributed systems is a plus.

Responsibilities

  • Design, deploy, and maintain large-scale GPU clusters for training and inference workloads.
  • Build automation pipelines for provisioning, scaling, and monitoring compute resources across Slurm and Kubernetes.
  • Develop observability, alerting, and auto-healing systems for high-availability GPU workloads.
  • Collaborate with ML, networking, and platform teams to optimise resource scheduling, GPU utilisation, and data flow.
  • Implement infrastructure-as-code, CI/CD pipelines, and reliability standards across thousands of nodes.
  • Diagnose performance bottlenecks and drive continuous improvements in reliability, latency, and throughput.

Skills

SRE/DevOps experience
Kubernetes
Slurm
Linux systems
Python/Go/Bash
Observability stacks
AI/ML infra experience
Reliability engineering

Tools

Prometheus
Grafana
Loki
CI/CD tooling

Job description

Hamilton Barnes Associates Limited is seeking an experienced SRE/Infrastructure Engineer to help build a seed-stage AI infrastructure company with large-scale GPU clusters for training and inference. You design, deploy, and maintain compute environments and ensure reliability across Slurm and Kubernetes.

You will implement automation, observability, and IaC practices, collaborating with ML and platform teams to optimize GPU utilization, data flow, and latency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE, AI Infrastructure & GPU Fleet Reliability
Senior SRE, AI Infrastructure & GPU Fleet Reliability

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Senior SRE - AI Infrastructure
Senior SRE - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
IPO Equity
10% comapny bonus
401K 4% match
Staff AI Infrastructure Engineer — Orchestration & Inference
Staff AI Infrastructure Engineer — Orchestration & Inference

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership
Senior AI Network Architect for Large-Scale GPU Clusters
Senior AI Network Architect for Large-Scale GPU Clusters

Hamilton Barnes Associates Limited • United States

On-site
USD 220,000 - 350,000
Annual bonus
Equity opportunities
Flexible working arrangements
+1
Senior AI Storage Engineer - Remote GPU HPC Infra
Senior AI Storage Engineer - Remote GPU HPC Infra

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Company bonus
Remote working options and allowance
Senior SRE — Global HPC/AI Infra with Equity
Senior SRE — Global HPC/AI Infra with Equity

Socket.dev • North Carolina

On-site
USD 152,000 - 288,000
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior AI Cloud SRE — HPC & GPU Infra
Senior AI Cloud SRE — HPC & GPU Infra

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+4
Senior SRE: AI/GPU Scale & Automation
Senior SRE: AI/GPU Scale & Automation

Nscale • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work expectations
Senior AI Cloud SRE — HPC & GPU Infrastructure
Senior AI Cloud SRE — HPC & GPU Infrastructure

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000