Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited

San Francisco (CA)

On-site

USD 270,000 - 330,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity

Job summary

A stealth-mode startup in San Francisco seeks a Platform Engineer/Senior Site Reliability Engineer to manage their AI and cloud platform. You will design and maintain large-scale GPU clusters, create automation pipelines, and enhance system reliability. Ideal candidates have over 7 years of experience in SRE or DevOps, strong skills in Kubernetes, Linux systems, and automation using Python or Go. This role offers a gross salary of $300,000 annually along with equity options.

Qualifications

  • 7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles.
  • Deep knowledge of Linux systems, networking, and GPU infrastructure.
  • Strong hands-on experience with Kubernetes and Slurm.

Responsibilities

  • Design, deploy, and maintain large-scale GPU clusters.
  • Build automation pipelines for provisioning, scaling, and monitoring resources.
  • Develop observability and auto-healing systems for high-availability workloads.

Skills

SRE experience
Kubernetes
Slurm
Linux systems
Python
Go
Bash
Prometheus
Grafana
Loki

Job description

Join a stealth-mode startup building out their AI and cloud platform, powered by thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring seamless orchestration across environments managed by Slurm, Kubernetes, or direct SSH access. As well as supporting their extremely exciting new products coming to the market!

This is a rare opportunity to work at the intersection of AI infrastructure and AI, shaping the operational backbone of one of the largest GPU clusters in private deployment.

If you want to build and operate infrastructure for frontier AI workloads, automate systems at petascale, and be part of a founding engineering team, this is the place to do it. Get in touch and apply today!

Responsibilities:
  • Design, deploy, and maintain large-scale GPU clusters (H100/H200/B200) for training and inference workloads.
  • Build automation pipelines for provisioning, scaling, and monitoring compute resources across Slurm and Kubernetes environments.
  • Develop observability, alerting, and auto-healing systems for high-availability GPU workloads.
  • Collaborate with ML, networking, and platform teams to optimise resource scheduling, GPU utilisation, and data flow.
  • Implement infrastructure-as-code, CI/CD pipelines, and reliability standards across thousands of nodes.
  • Diagnose performance bottlenecks and drive continuous improvements in reliability, latency, and throughput.
Skills / Must Have:
  • 7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting large-scale compute environments.
  • Strong hands-on experience with Kubernetes and Slurm for cluster orchestration and workload management.
  • Deep knowledge of Linux systems, networking, and GPU infrastructure (NVIDIA H100/H200/B200 preferred).
  • Proficiency in Python, Go, or Bash for automation, tooling, and performance tuning.
  • Experience with observability stacks (Prometheus, Grafana, Loki) and incident response frameworks.
  • Familiarity with high-performance computing (HPC) or AI/ML training infrastructure at scale.
  • Background in reliability engineering, distributed systems, or hardware acceleration environments is a strong plus.
Salary & Benefits:
  • $300,000 gross per year
  • Equity
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Infrastructure Product Engineer - AI Infrastructure
Infrastructure Product Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 300,000 - 350,000
Equity
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Platform Engineer (GPU)
Platform Engineer (GPU)

Vero • United States

On-site
USD 136,000 - 160,000
Medical, dental, and vision insurance
Equity Scheme
401(k) with employer match
+3
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • San Francisco (CA)

On-site
USD 175,000 - 250,000
Senior GPU Infrastructure Engineer - AI Infrastructure
Senior GPU Infrastructure Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus