Remote Forward Deployed SRE Engineer for GPU Clusters

andromeda hill

San Francisco (CA)

Hybrid

USD 160,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Meaningful equity
Healthcare
Dental
Vision
401(k)
Unlimited PTO

Job summary

Andromeda is hiring a Forward Deployed Engineer - SRE to embed with customer teams running large-scale GPU training and inference. You will onboard customers, choose orchestration (Slurm, Kubernetes, or SSH), and ensure reliability across clusters with automation and monitoring.

You should have hands-on production GPU experience, deep Linux, and IaC skills, plus the ability to communicate complex issues to customer leaders and drive product improvements.

Qualifications

  • Hands-on experience operating GPU clusters in production.
  • Production experience with high-speed fabrics (InfiniBand, RoCE, NVLink).
  • Deep understanding of distributed training workflows and performance bottlenecks.
  • Expert-level Linux experience including kernel tuning and driver management.
  • Proficiency with Kubernetes in production for GPU workloads.
  • Strong software engineering skills (Python/Go/Bash).
  • Infrastructure-as-Code tooling experience (Terraform, Helm, Ansible).
  • Experience building monitoring/alerting for GPU telemetry.
  • Ability to articulate complex technical tradeoffs to customer teams.
  • Proven incident response leadership for complex distributed systems.

Responsibilities

  • Serve as primary technical contact for large-scale training/inference workloads and onboarding.
  • Diagnose real failures in customer environments; reproduce, fix, and document.
  • Profile and improve distributed training performance and reduce idle GPU time.
  • Own reliability outcomes for deployed customer accounts.
  • Maintain health of high-speed interconnects (InfiniBand, RoCE, NVLink).
  • Develop deep visibility into GPU utilization, memory pressure, and hardware health.
  • Automate repeated deployments: provisioning, health checks, preflight validation, self-healing.
  • Lead incident response including blameless postmortems and systemic fixes.
  • Contribute to roadmap by surfacing customer-visible signals from field.

Skills

GPU clusters prod
GPU fabrics
Distributed training
Linux expert
Kubernetes w/GPU
Python/Go/Bash
IaC (Terraform/Helm/Ansible)
GPU telemetry monitoring
Incident response
Architecture communication

Tools

Slurm
SSH

Job description

Andromeda is hiring a Forward Deployed Engineer - SRE to embed with customer teams running large-scale GPU training and inference. You will onboard customers, choose orchestration (Slurm, Kubernetes, or SSH), and ensure reliability across clusters with automation and monitoring.

You should have hands-on production GPU experience, deep Linux, and IaC skills, plus the ability to communicate complex issues to customer leaders and drive product improvements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Forward-Deployed SRE for AI Infra
Remote Forward-Deployed SRE for AI Infra

Andromeda Cluster • United States

Hybrid
USD 180,000 - 240,000
Remote FDE-SRE: GPU Cluster Reliability & Onboarding
Remote FDE-SRE: GPU Cluster Reliability & Onboarding

Andromeda Cluster, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Competitive equity package
Healthcare, dental, vision
401(k) plan
+1
Remote GPU Infra SRE for Large-Scale AI Training
Remote GPU Infra SRE for Large-Scale AI Training

andromeda hill • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

Andromeda Cluster, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Competitive equity package
Healthcare, dental, vision
401(k) plan
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Andromeda • San Francisco (CA)

On-site
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

Andromeda Cluster • United States

Hybrid
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

andromeda hill • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

andromeda hill • San Francisco (CA)

Hybrid
USD 160,000 - 210,000
Meaningful equity
Healthcare
Dental
+3
Member of the Technical Staff - Systems
Member of the Technical Staff - Systems

Andromeda • San Francisco (CA)

Hybrid
USD 190,000 - 260,000
Customer Reliability Engineer
Customer Reliability Engineer

Andromeda Cluster • San Francisco (CA)

On-site
USD 120,000 - 160,000