Remote Forward-Deployed SRE for AI Infra

Andromeda Cluster

United States

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Andromeda Cluster is seeking a Forward Deployed Engineer - SRE to work inside customer environments, optimizing large-scale GPU training and inference pipelines. You will onboard teams, tune runs, and debug failures, owning reliability of the platform and high-performance interconnects.

You will collaborate with customers to diagnose failures, reproduce issues, and ship fixes that improve the platform. This role blends hands-on ops with customer-facing engineering in a high-growth AI infra

Qualifications

  • Hands-on experience operating GPU clusters in production.
  • Production experience with InfiniBand, RoCE, or NVLink fabrics.
  • Understanding GPU memory hierarchies, ECC behavior, and failure modes from direct experience.
  • Production-grade Kubernetes with GPU workloads experience.
  • Strong programming in Python, Go, or Bash.
  • Infra-as-Code (Terraform, Helm, Ansible).
  • Experience leading incident response for distributed systems.
  • Ability to explain findings to a customer team without condescension.

Responsibilities

  • Own onboarding end to end for teams running large-scale training/inference workloads.
  • Diagnose real failures in customer environments: NCCL timeouts, I/O stalls, degraded links.
  • Profile and improve distributed training performance on live workloads.
  • Own reliability outcomes for the accounts you’re deployed on.
  • Ensure health of high-speed interconnects (InfiniBand, RoCE, NVLink).
  • Build monitoring for GPU telemetry and dashboards.
  • Turn repeated deployments into automation: provisioning, health checks, preflight validation.
  • Lead incident response and postmortem with systemic fixes.

Skills

GPU clusters
Kubernetes
Slurm
NVIDIA drivers
CUDA toolkit
Python
Go
Terraform

Tools

NVIDIA driver management
Container runtimes
DCGM / nvidia-smi

Job description

Andromeda Cluster is seeking a Forward Deployed Engineer - SRE to work inside customer environments, optimizing large-scale GPU training and inference pipelines. You will onboard teams, tune runs, and debug failures, owning reliability of the platform and high-performance interconnects.

You will collaborate with customers to diagnose failures, reproduce issues, and ship fixes that improve the platform. This role blends hands-on ops with customer-facing engineering in a high-growth AI infra

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote FDE-SRE: GPU Cluster Reliability & Onboarding
Remote FDE-SRE: GPU Cluster Reliability & Onboarding

Andromeda Cluster, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Competitive equity package
Healthcare, dental, vision
401(k) plan
+1
Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

Andromeda Cluster, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Competitive equity package
Healthcare, dental, vision
401(k) plan
+1
Remote AI Infrastructure SRE — Kubernetes & Reliability
Remote AI Infrastructure SRE — Kubernetes & Reliability

Andromeda • San Francisco (CA)

On-site
USD 120,000 - 160,000
Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

Andromeda Cluster • United States

Hybrid
USD 180,000 - 240,000
Customer Reliability Engineer
Customer Reliability Engineer

Andromeda Cluster • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Andromeda • San Francisco (CA)

On-site
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Remote AI Inference & Post-Training Specialist
Remote AI Inference & Post-Training Specialist

Togetherai • San Francisco (CA)

On-site
USD 270,000 - 300,000
Startup equity
Health insurance
Remote work flexibility
Remote AI Cloud Forward-Deployed Engineer
Remote AI Cloud Forward-Deployed Engineer

3M HEALTHCARE • San Francisco (CA)

On-site
USD 195,000 - 239,000
Staff Forward Deployed Engineer, AI/ML
Staff Forward Deployed Engineer, AI/ML

digitalocean98 • United States

Hybrid
USD 220,000 - 239,000
Bonus
Equity
Hybrid work model
Staff AI/ML Forward-Deployed Engineer, Production Lead
Staff AI/ML Forward-Deployed Engineer, Production Lead

digitalocean98 • United States

Hybrid
USD 220,000 - 239,000
Bonus
Equity
Hybrid work model