Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Andromeda Cluster is seeking a Senior Site Reliability Engineer to design, operate and debug large-scale GPU infrastructure used for distributed training and inference, working directly with customers pushing the limits of modern AI systems.

You will own GPU cluster architecture, optimize performance, ensure reliability, and build automation and observability tooling. The role requires hands-on experience with GPU hardware, Linux internals, Kubernetes and high-speed interconnects.

Qualifications

  • Hands-on experience operating large-scale GPU clusters.
  • Deep knowledge of GPU memory hierarchies, ECC behavior, and failure modes.
  • Production networking with InfiniBand, RoCE, or NVLink in distributed training.
  • Proficiency with Linux kernel tuning, drivers, and performance profiling.
  • Experience with Kubernetes for GPU workloads and device plugins.

Responsibilities

  • GPU Cluster Architecture: design multi-provider, multi-region GPU clusters for large-scale training.
  • Customer Technical Partnership: onboard and troubleshoot customers running large-scale training workloads.
  • Reliability & Performance Engineering: define SLOs and budgets for GPU infrastructure.
  • Networking & Fabric Health: diagnose and resolve high-speed interconnect issues.
  • Observability: build visibility into GPU utilization and training performance.
  • Automation & Tooling: build automation for provisioning, health checks, and scheduling.
  • Incident Leadership: lead postmortems and drive systemic fixes.

Skills

GPU Systems Expertise
High-Performance Networking
Distributed Training & ML Frameworks
Linux & Systems Internals
Kubernetes & Orchestration
Automation & Software Engineering
Observability & Monitoring
Incident Management

Tools

NVIDIA DCGM
nvidia-smi monitoring

Job description

Andromeda Cluster is seeking a Senior Site Reliability Engineer to design, operate and debug large-scale GPU infrastructure used for distributed training and inference, working directly with customers pushing the limits of modern AI systems.

You will own GPU cluster architecture, optimize performance, ensure reliability, and build automation and observability tooling. The role requires hands-on experience with GPU hardware, Linux internals, Kubernetes and high-speed interconnects.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Andromeda • San Francisco (CA)

On-site
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Embedded SRE for AI Training Clusters - Remote + Equity
Embedded SRE for AI Training Clusters - Remote + Equity

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Health insurance
Equity
Unlimited PTO
+1
Senior AI Infra Engineer - Scalable GPU Clusters
Senior AI Infra Engineer - Scalable GPU Clusters

NVIDIA • Washington

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior AI Infra SRE: GPU Clusters & High-Perf Networking
Senior AI Infra SRE: GPU Clusters & High-Perf Networking

Andromeda • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Senior Full-Stack Engineer, AI Infra for GPU Clusters
Senior Full-Stack Engineer, AI Infra for GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000