Forward Deployed Engineer - SRE

Andromeda Cluster

United States

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Andromeda Cluster is seeking a Forward Deployed Engineer - SRE to work inside customer environments, optimizing large-scale GPU training and inference pipelines. You will onboard teams, tune runs, and debug failures, owning reliability of the platform and high-performance interconnects.

You will collaborate with customers to diagnose failures, reproduce issues, and ship fixes that improve the platform. This role blends hands-on ops with customer-facing engineering in a high-growth AI infra

Qualifications

  • Hands-on experience operating GPU clusters in production.
  • Production experience with InfiniBand, RoCE, or NVLink fabrics.
  • Understanding GPU memory hierarchies, ECC behavior, and failure modes from direct experience.
  • Production-grade Kubernetes with GPU workloads experience.
  • Strong programming in Python, Go, or Bash.
  • Infra-as-Code (Terraform, Helm, Ansible).
  • Experience leading incident response for distributed systems.
  • Ability to explain findings to a customer team without condescension.

Responsibilities

  • Own onboarding end to end for teams running large-scale training/inference workloads.
  • Diagnose real failures in customer environments: NCCL timeouts, I/O stalls, degraded links.
  • Profile and improve distributed training performance on live workloads.
  • Own reliability outcomes for the accounts you’re deployed on.
  • Ensure health of high-speed interconnects (InfiniBand, RoCE, NVLink).
  • Build monitoring for GPU telemetry and dashboards.
  • Turn repeated deployments into automation: provisioning, health checks, preflight validation.
  • Lead incident response and postmortem with systemic fixes.

Skills

GPU clusters
Kubernetes
Slurm
NVIDIA drivers
CUDA toolkit
Python
Go
Terraform

Tools

NVIDIA driver management
Container runtimes
DCGM / nvidia-smi

Job description

Forward Deployed Engineer - SRE
Location: North America Remote/SF-Hybrid · Full-Time

About Andromeda

Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.

We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible.

Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.

Our long-term vision is to build the liquidity layer for global AI compute. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.

The Role

This is not a generalist SRE role, and it is not a support role. You will embed directly with the teams running large-scale training and inference on our clusters. You are responsible for onboarding them, tuning their jobs, and debugging their failures alongside them, while owning the infrastructure and automation that makes those clusters reliable in the first place.

Forward deployed means you spend real time inside customer environments: reading their training code, sitting in their Slack channels, watching their runs, and shipping fixes that land in our platform. When a multi-hundred-GPU run stalls, you are the person who figures out whether it’s the fabric, the driver, the scheduler, or their dataloader, then you make sure it can’t happen the same way twice.

We’re looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric → kernel → framework. Equally important: you can explain what you found to someone else’s engineering team without condescension, and turn that conversation into a product improvement.

What You’ll Do

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads. Own onboarding end to end; environment setup, orchestration choice (Slurm, Kubernetes, or direct SSH), storage layout, first successful run at scale. You will continue to stay engaged as their workloads grow.

  • Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns, container and driver mismatches. Read their code when you need to. Reproduce, isolate, fix, and write it down.

  • Profile and improve distributed training performance on live workloads. Improving MFU, cutting idle GPU time, and reducing time-to-first-successful-run for new deployments.

  • Own reliability outcomes for the accounts you’re deployed on.

  • Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) that underpin distributed training. Diagnose and resolve fabric-level issues that degrade collective operations.

  • Build deep visibility into GPU utilization, memory pressure, interconnect throughput, job performance, and hardware health.

  • Turn every repeated deployment problem into automation: cluster provisioning, GPU health checks and burn-in, preflight validation, self-healing, firmware/driver lifecycle management, and reusable reference configurations for common training and serving stacks.

  • Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks. Own the customer-facing communication during the incident and the blameless postmortem and systemic fix after it.

  • You will see our rough edges before anyone else does. Bring that signal back to influence the roadmap, file the hard bugs, and build the missing pieces yourself when that’s the fastest path.

What We’re Looking For

  • Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent). You understand GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes from direct experience.

  • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training. You can diagnose why an all-reduce is slow, identify a degraded link in a fat-tree topology, and reason about congestion control at scale.

  • Working knowledge of how large training and inference jobs actually run. You don’t need to design models, but you need to understand what’s happening at the systems level when a large run stalls.

  • Expert-level Linux experience: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at the syscall and hardware level.

  • Strong experience running Kubernetes in production with GPU workloads. Experience with device plugins, topology-aware scheduling, multi-cluster, custom operators. Experience with Slurm or other HPC schedulers is equally valued.

  • Strong engineering skills in Python, Go, or Bash. You build production-grade tools and services, not just scripts.

  • Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent).

  • Hands-on experience building monitoring and alerting for GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics) integrated into actionable dashboards.

  • You can go deep on architecture with a customer’s infra team and clearly articulate tradeoffs to their leadership. You’re comfortable being the only one in the room who knows the answer, and equally comfortable saying you don’t yet.

  • Proven track record leading incident response for complex distributed systems.

Strong Candidates May Have

  • Experience with high-performance parallel file systems (VAST, WEKA, Lustre, GPFS) and the checkpoint I/O and data-loading bottlenecks that come with large training runs.

  • Time spent embedded with external engineering teams, i.e solutions architecture, professional services, deployed SRE, or technical account ownership at an infrastructure company.

  • Experience operating production inference. Hands on experience with autoscaling, batching, KV cache behavior, cold starts, and multi-tenant GPU sharing.

  • Contributions to relevant OSS projects, or benchmarks, postmortems, and deep-dives you’ve published.

  • Experience working across heterogeneous providers and regions rather than a single hyperscaler.

What Success Looks Like

By the end of your first year:

  • You know each of your accounts’ actual technical goals including what they’re training, what their scaling roadmap looks like over the next two quarters, what their real constraints are (budget, deadline, headcount, data), and you’ve written that down somewhere the rest of us can read it.

  • You are the person your accounts’ engineers message first, before they file a ticket, because you’ve earned it.

  • Their reliability and throughput numbers are visibly better than at onboarding, and you can point to the specific changes that did it.

  • Recurring problems you found in the field exist as automation, preflight checks, or documentation, not as tribal knowledge in your head.

  • You’ve advocated internally for at least one roadmap change on behalf of a strategic customer, and it shipped.

Why You’ll Love It Here

  • High-growth environment: Get in early at a company at the center of the AI infrastructure boom

  • Ownership: First FDE for the solutions engineering team, you’ll get to build this function from the ground up

  • Competitive compensation: + meaningful equity

  • Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO

Andromeda Cluster is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

Andromeda Cluster, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Competitive equity package
Healthcare, dental, vision
401(k) plan
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Andromeda • San Francisco (CA)

On-site
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Solutions Architect
Solutions Architect

Andromeda • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Equity
Healthcare, dental, and vision
401(k)
+1
Technical Program Manager
Technical Program Manager

Andromeda • San Francisco (CA)

Hybrid
USD 140,000 - 200,000
Member of the Technical Staff - Systems
Member of the Technical Staff - Systems

Andromeda Cluster, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Competitive compensation
Equity
Healthcare
+4
Revenue Operations Lead
Revenue Operations Lead

Andromeda • United States

Hybrid
USD 120,000 - 190,000
Equity
Healthcare, dental, and vision
401(k) and unlimited PTO
Compute Procurement Lead
Compute Procurement Lead

Andromeda Cluster, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Healthcare, dental, and vision
401(k)
Unlimited PTO
People Operations Generalist
People Operations Generalist

Andromeda • San Francisco (CA)

Hybrid
USD 120,000 - 170,000
Healthcare
401(k)
Unlimited PTO
+1
People Operations Generalist
People Operations Generalist

Andromeda Cluster • United States

Hybrid
USD 90,000 - 120,000
Healthcare
Dental
Vision
+2
Strategic Compute Finance Lead
Strategic Compute Finance Lead

Andromeda • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive compensation with equity
Comprehensive benefits including healthcare
Unlimited PTO