Remote GPU Infra SRE for Large-Scale AI Training

andromeda hill

San Francisco (CA)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Andromeda is seeking a Senior Site Reliability Engineer to design, operate, and optimize large-scale GPU infrastructure used for distributed training and inference. You will work directly with customers to push the limits of AI systems and ensure reliable capacity across multiple regions.

You will own GPU cluster architecture, performance, and reliability engineering, with emphasis on network fabrics, observability, automation, and incident leadership in production environments.

Qualifications

  • Deep, hands-on experience operating large-scale GPU clusters.
  • Experience with GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes.

Responsibilities

  • GPU Cluster Architecture: design and evolve multi-provider, multi-region GPU clusters.
  • Customer Technical Partnership: onboard and troubleshoot for customers running large-scale training workloads.
  • Reliability & Performance Engineering: define SLOs and budgets for GPU infrastructure.
  • Networking & Fabric Health: ensure health of InfiniBand, RoCE, NVLink fabrics.

Skills

GPU systems
High-performance networking
Distributed training
Linux internals
Kubernetes
Automation
Observability
Incident management

Tools

CUDA toolkit
NCCL
InfiniBand

Job description

Andromeda is seeking a Senior Site Reliability Engineer to design, operate, and optimize large-scale GPU infrastructure used for distributed training and inference. You will work directly with customers to push the limits of AI systems and ensure reliable capacity across multiple regions.

You will own GPU cluster architecture, performance, and reliability engineering, with emphasis on network fabrics, observability, automation, and incident leadership in production environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Andromeda • San Francisco (CA)

On-site
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Remote Forward Deployed SRE Engineer for GPU Clusters
Remote Forward Deployed SRE Engineer for GPU Clusters

andromeda hill • San Francisco (CA)

Hybrid
USD 160,000 - 210,000
Meaningful equity
Healthcare
Dental
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

andromeda hill • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Remote Forward-Deployed SRE for AI Infra
Remote Forward-Deployed SRE for AI Infra

Andromeda Cluster • United States

Hybrid
USD 180,000 - 240,000
Remote FDE-SRE: GPU Cluster Reliability & Onboarding
Remote FDE-SRE: GPU Cluster Reliability & Onboarding

Andromeda Cluster, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Competitive equity package
Healthcare, dental, vision
401(k) plan
+1
Senior GPU Infra Architect for Scalable AI Compute
Senior GPU Infra Architect for Scalable AI Compute

AI Chopping Block • Costa Mesa (CA), Northern (KY)

Hybrid
USD 166,000 - 220,000
Senior AI Infra SRE: GPU Clusters & High-Perf Networking
Senior AI Infra SRE: GPU Clusters & High-Perf Networking

Andromeda • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Remote AI Infrastructure SRE — Kubernetes & Reliability
Remote AI Infrastructure SRE — Kubernetes & Reliability

Andromeda • San Francisco (CA)

On-site
USD 120,000 - 160,000
Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

Andromeda Cluster, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Competitive equity package
Healthcare, dental, vision
401(k) plan
+1
Senior SRE — AI GPU Infra Architect (Multi-Cloud)
Senior SRE — AI GPU Infra Architect (Multi-Cloud)

lumalabs-ai • San Francisco (CA)

On-site
USD 170,000 - 290,000