Senior AI Infra SRE: GPU Clusters & High-Perf Networking
Andromeda
San Francisco (CA)
Hybrid
USD 150,000 - 200,000
Full time
14 days+
Application generator
An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Get past ATS filters
Benefits offered by this job
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Job summary
A leading AI infrastructure company is looking for a Senior Site Reliability Engineer to design and operate large-scale GPU clusters. In this role, you will work closely with clients to troubleshoot and optimize AI infrastructure. The ideal candidate has extensive experience with GPU systems, high-performance networking, and Linux internals. You’ll have significant influence over technical direction and help define operations for scalable AI compute, all while ensuring system reliability and efficiency.
Expert-level Linux knowledge: kernel tuning, driver management.
Strong engineering skills in Python, Go, or Bash.
Proven track record leading incident response.
Responsibilities
Design and evolve multi-provider, multi-region GPU compute clusters.
Serve as the primary technical point of contact for customers.
Define SLOs and error budgets for GPU infrastructure.
Build deep visibility into GPU utilization and performance.
Skills
GPU Systems Expertise
High-Performance Networking
Distributed Training & ML Frameworks
Linux & Systems Internals
Kubernetes & Orchestration
Automation & Software Engineering
Observability & Monitoring
Incident Management
Tools
Kubernetes
CUDA
Terraform
Python
Bash
Prometheus
Grafana
Job description
A leading AI infrastructure company is looking for a Senior Site Reliability Engineer to design and operate large-scale GPU clusters. In this role, you will work closely with clients to troubleshoot and optimize AI infrastructure. The ideal candidate has extensive experience with GPU systems, high-performance networking, and Linux internals. You’ll have significant influence over technical direction and help define operations for scalable AI compute, all while ensuring system reliability and efficiency.