Senior AI Infra SRE: GPU Clusters & High-Perf Networking

Andromeda

San Francisco (CA)

Hybrid

USD 150,000 - 200,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure

Job summary

A leading AI infrastructure company is looking for a Senior Site Reliability Engineer to design and operate large-scale GPU clusters. In this role, you will work closely with clients to troubleshoot and optimize AI infrastructure. The ideal candidate has extensive experience with GPU systems, high-performance networking, and Linux internals. You’ll have significant influence over technical direction and help define operations for scalable AI compute, all while ensuring system reliability and efficiency.

Qualifications

  • Deep, hands-on experience operating large-scale GPU clusters.
  • Expert-level Linux knowledge: kernel tuning, driver management.
  • Strong engineering skills in Python, Go, or Bash.
  • Proven track record leading incident response.

Responsibilities

  • Design and evolve multi-provider, multi-region GPU compute clusters.
  • Serve as the primary technical point of contact for customers.
  • Define SLOs and error budgets for GPU infrastructure.
  • Build deep visibility into GPU utilization and performance.

Skills

GPU Systems Expertise
High-Performance Networking
Distributed Training & ML Frameworks
Linux & Systems Internals
Kubernetes & Orchestration
Automation & Software Engineering
Observability & Monitoring
Incident Management

Tools

Kubernetes
CUDA
Terraform
Python
Bash
Prometheus
Grafana

Job description

A leading AI infrastructure company is looking for a Senior Site Reliability Engineer to design and operate large-scale GPU clusters. In this role, you will work closely with clients to troubleshoot and optimize AI infrastructure. The ideal candidate has extensive experience with GPU systems, high-performance networking, and Linux internals. You’ll have significant influence over technical direction and help define operations for scalable AI compute, all while ensuring system reliability and efficiency.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI GPU Infra SRE - Scale, Automation & Equity
Senior AI GPU Infra SRE - Scale, Automation & Equity

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior AI GPU Infra Engineer — Performance & Scale
Senior AI GPU Infra Engineer — Performance & Scale

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior AI Infra Engineer: GPU Clusters & Kubernetes
Senior AI Infra Engineer: GPU Clusters & Kubernetes

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Senior SRE: AI Infra on-Site in SF, GPU & Cloud
Senior SRE: AI Infra on-Site in SF, GPU & Cloud

The Recruiting Guy • Arlington (VA)

On-site
USD 175,000 - 250,000
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead
Senior GPU Infra Engineer: AI Clusters & OpenStack Lead

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 120,000 - 160,000
Potential equity/bonus
Senior AI Infra Networking Engineer | High-Perf GPU Cloud
Senior AI Infra Networking Engineer | High-Perf GPU Cloud

Nscale • United States

Remote
USD 100,000 - 200,000
Senior Network Engineer: AI Data Centers & GPU Cloud
Senior Network Engineer: AI Data Centers & GPU Cloud

Hamilton Barnes ? • United States

Hybrid
USD 120,000 - 190,000
Medical insurance
Vision insurance
401(k)
+3
Senior AI Infra SRE — GPU Cloud Reliability Leader
Senior AI Infra SRE — GPU Cloud Reliability Leader

deCircle • San Francisco (CA)

On-site
USD 120,000 - 150,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000