Customer Reliability Engineer

Andromeda Cluster

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A cutting-edge AI infrastructure company is seeking a Site Reliability Engineer to manage Kubernetes clusters and improve the reliability of critical systems. The ideal candidate will have 5+ years of experience in SRE or DevOps, strong Linux and Kubernetes expertise, and skills in automation and Infrastructure-as-Code. This role offers the opportunity to shape the future of scalable AI infrastructure, working closely with both customers and technical teams in a dynamic environment.

Qualifications

  • 5+ years experience in SRE, DevOps, or infrastructure engineering roles.
  • Strong Linux systems and networking fundamentals.
  • Deep experience with Kubernetes and container orchestration at scale.

Responsibilities

  • Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers.
  • Build automation and tooling to streamline cluster deployments and integrations.
  • Collaborate with engineering and product teams to plan and deliver infrastructure for new services.

Skills

SRE experience
Linux systems knowledge
Kubernetes expertise
Infrastructure-as-Code proficiency
Scripting skills in Python, Go, or Bash
Automation skills
Experience with observability tools

Tools

Terraform
Ansible
Prometheus
Grafana

Job description

Site Reliability Engineer - AI Infrastructure

Location: Global Remote / San Francisco · Full-Time

About Andromeda

Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.

We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible.

Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.

Our long-term vision is to build the liquidity layer for global AI compute — a marketplace that moves the infrastructure and workloads powering AGI not dissimilar to the flows of capital in the world's financial markets.

We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.

What You’ll Do
  • Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers.

  • Build automation and tooling to streamline cluster deployments and integrations.

  • Debug customer issues across networking, storage, scheduling, and system layers.

  • Improve reliability and scalability of both training and inference infrastructure.

  • Design and implement monitoring, alerting, and observability for critical systems.

  • Collaborate with engineering and product teams to plan and deliver infrastructure for new services.

  • Participate in on-call and incident response, leading postmortems and reliability improvements.

What We’re Looking For
  • 5+ years experience in SRE, DevOps, or infrastructure engineering roles.

  • Strong Linux systems and networking fundamentals.

  • Deep experience with Kubernetes and container orchestration at scale.

  • Proficiency with Infrastructure-as-Code (Terraform, Helm, Ansible, etc.).

  • Strong automation and scripting skills (Python, Go, or Bash).

  • Experience with observability stacks (Prometheus, Grafana, Loki, Datadog, etc.).

  • Track record of operating production systems and leading incident response.

Nice to Have
  • Exposure to ML/AI infrastructure or GPU-based systems (CUDA, Slurm, Triton, etc.).

  • Familiarity with high-performance networking (InfiniBand, NVLink) or distributed storage (VAST, Weka, Ceph).

  • Customer-facing support or consulting experience.

Why You’ll Love It Here

This is a builder’s role. You’ll have ownership and autonomy to shape how our systems run, working directly with customers and providers while building the foundation for reliable, scalable AI infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Andromeda • San Francisco (CA)

Remote
USD 120,000 - 160,000
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Ownership and autonomy in projects
Engagement with customers and providers
Opportunity to shape systems
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Andromeda • San Francisco (CA)

On-site
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Remote AI Infrastructure SRE — Kubernetes & Reliability
Remote AI Infrastructure SRE — Kubernetes & Reliability

Andromeda • San Francisco (CA)

On-site
USD 120,000 - 160,000
Forward Deployed Engineer - SRE
Forward Deployed Engineer - SRE

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Health insurance
Equity
Unlimited PTO
+1
Member of the Business Staff - Compute Markets
Member of the Business Staff - Compute Markets

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 90,000 - 120,000
Competitive compensation
Meaningful equity
Comprehensive healthcare benefits
+1
Member of the Technical Staff - Systems
Member of the Technical Staff - Systems

Andromeda • San Francisco (CA)

Hybrid
USD 190,000 - 260,000
Solutions Architect
Solutions Architect

Andromeda • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Equity
Healthcare, dental, and vision
401(k)
+1
Infrastructure Manager
Infrastructure Manager

The Resume Database • San Francisco (CA)

On-site
USD 100,000 - 130,000
Competitive compensation
Meaningful equity
Comprehensive benefits including healthcare
+1