Remote FDE-SRE: GPU Cluster Reliability & Onboarding

Andromeda Cluster, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive equity package
Healthcare, dental, vision
401(k) plan
Unlimited PTO

Job summary

Andromeda Cluster, Inc. is hiring a Forward Deployed Engineer - SRE to join our remote-first team serving North America.

You will embed with customers running large-scale GPU training and inference, onboarding, tuning jobs, and debugging failures while owning the reliability of the platform. You will partner with customer teams inside their environments, diagnose complex issues, and drive they scale from day one.

Qualifications

  • Hands-on experience operating GPU clusters in production.
  • Experience with distributed training and high-performance networking.
  • Strong Linux/Unix troubleshooting and performance tuning.

Responsibilities

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads.
  • Own onboarding end-to-end: environment setup, orchestration choice (Slurm, Kubernetes, SSH), storage layout, first successful run at scale.
  • Diagnose real failures in customer environments: NCCL timeouts, stragglers, I/O stalls, degraded links.
  • Profile and improve distributed training performance and reduce idle GPU time.
  • Lead incident response for complex failures across hardware, networking, and ML frameworks.
  • Turn repeated deployment problems into automation and reusable configurations.

Skills

GPU clusters
Kubernetes
Linux
Incident response

Education

Bachelor's degree in CS or related field

Tools

Slurm
Terraform
DCGM/nvidia-smi

Job description

Andromeda Cluster, Inc. is hiring a Forward Deployed Engineer - SRE to join our remote-first team serving North America.

You will embed with customers running large-scale GPU training and inference, onboarding, tuning jobs, and debugging failures while owning the reliability of the platform. You will partner with customer teams inside their environments, diagnose complex issues, and drive they scale from day one.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Forward-Deployed SRE for AI Infra
Remote Forward-Deployed SRE for AI Infra

Andromeda Cluster • United States

Hybrid
USD 180,000 - 240,000
Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

Andromeda Cluster, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Competitive equity package
Healthcare, dental, vision
401(k) plan
+1
Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

Andromeda Cluster • United States

Hybrid
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Andromeda • San Francisco (CA)

On-site
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
SRE: AI GPU Infra for High-Security Clusters
SRE: AI GPU Infra for High-Security Clusters

United States Digital Space LLC • Washington, El Segundo (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, dental coverage
+1
Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Senior GPU Infra Reliability Engineer - Remote
Senior GPU Infra Reliability Engineer - Remote

Luma AI • United States

Remote
USD 180,000 - 240,000
Senior SRE & Automation Engineer — GPU Cloud Reliability
Senior SRE & Automation Engineer — GPU Cloud Reliability

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 230,000
Senior SRE - GPU Cloud Reliability & Automation
Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000
Senior SRE - Customer-Facing GPU Cloud Reliability
Senior SRE - Customer-Facing GPU Cloud Reliability

Bitdeer • San Jose (CA)

On-site
USD 180,000 - 260,000