Senior GPU HPC SRE — Remote, Stock Options

Hamilton Barnes Associates Limited

San Francisco (CA)

On-site

USD 225,000 - 275,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock options
Bonus
Remote working option and allowance

Job summary

Hamilton Barnes Associates Limited is seeking a Senior / Staff Site Reliability Engineer in San Francisco, California. This role focuses on supporting and scaling HPC and cloud environments, improving automation and reliability across distributed systems.

The ideal candidate will possess deep experience in Site Reliability Engineering and strong Linux expertise, along with automation skills in Python, Go, or Bash. The company offers competitive benefits including stock options and a remote working allowance.

Qualifications

  • Deep experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
  • Strong experience in large-scale distributed computing environments.
  • Deep Linux expertise (Ubuntu/Debian preferred).
  • Strong scripting and automation skills.
  • Experience with public cloud platforms or modern GPU cloud providers.

Responsibilities

  • Ensure reliability and performance of HPC and cloud infrastructure.
  • Design, build, and maintain automation and monitoring for GPU clusters.
  • Collaborate with teams on infrastructure systems.
  • Improve CI/CD pipelines and operational tooling.
  • Diagnose performance bottlenecks across distributed systems.
  • Support Slurm-based GPU cluster environments.

Skills

Site Reliability Engineering
DevOps
Infrastructure Engineering
Python
Go
Bash
Linux
Networking Fundamentals
Infrastructure-as-Code
Slurm

Tools

Terraform
Ansible

Job description

Hamilton Barnes Associates Limited is seeking a Senior / Staff Site Reliability Engineer in San Francisco, California. This role focuses on supporting and scaling HPC and cloud environments, improving automation and reliability across distributed systems.

The ideal candidate will possess deep experience in Site Reliability Engineering and strong Linux expertise, along with automation skills in Python, Go, or Bash. The company offers competitive benefits including stock options and a remote working allowance.

Get your free, confidential resume review.
or drag and drop your file here.