Remote AI Infrastructure SRE — Kubernetes & Reliability

Andromeda

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Andromeda is seeking a Site Reliability Engineer to work on AI infrastructure. This role involves provisioning Kubernetes clusters and improving reliability and scalability of systems. You will have the opportunity to automate processes and directly engage with customers.

Ideal candidates will have 5+ years in SRE or DevOps, strong networking fundamentals, and deep experience with Kubernetes. This position offers a unique opportunity to shape infrastructure for scalable AI.

Qualifications

  • 5+ years experience in SRE, DevOps, or infrastructure engineering roles.
  • Strong Linux systems and networking fundamentals are required.
  • Experience with Kubernetes and container orchestration at scale.

Responsibilities

  • Provision, configure, and operate Kubernetes-based clusters for customers.
  • Build automation and tooling to streamline cluster deployments.
  • Debug customer issues across networking, storage, scheduling, and system layers.
  • Improve infrastructure reliability and scalability.

Skills

SRE, DevOps, or infrastructure engineering experience
Linux systems and networking fundamentals
Kubernetes and container orchestration
Infrastructure-as-Code proficiency
Automation and scripting skills

Tools

Terraform
Python
Prometheus
Grafana

Job description

Andromeda is seeking a Site Reliability Engineer to work on AI infrastructure. This role involves provisioning Kubernetes clusters and improving reliability and scalability of systems. You will have the opportunity to automate processes and directly engage with customers.

Ideal candidates will have 5+ years in SRE or DevOps, strong networking fundamentals, and deep experience with Kubernetes. This position offers a unique opportunity to shape infrastructure for scalable AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Customer Reliability Engineer
Customer Reliability Engineer

Andromeda Cluster • San Francisco (CA)

On-site
USD 120,000 - 160,000
Global Remote SRE for AI Infrastructure & Kubernetes
Global Remote SRE for AI Infrastructure & Kubernetes

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Staff SRE Engineer: AI-Driven Platform Reliability
Staff SRE Engineer: AI-Driven Platform Reliability

AI Chopping Block • Costa Mesa (CA), Northern (KY)

Hybrid
USD 191,000 - 253,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
SRE: AI/ML Infra on Kubernetes, AWS & Terraform
SRE: AI/ML Infra on Kubernetes, AWS & Terraform

Deepgram • United States

Hybrid
USD 120,000 - 150,000
AI Reliability Engineer (AI SRE)
AI Reliability Engineer (AI SRE)

DeWinter Group • Campbell (CA)

Remote
AI Infra SRE Lead — Kubernetes, CI/CD & Observability
AI Infra SRE Lead — Kubernetes, CI/CD & Observability

The Consensus • United States

On-site
USD 263,000 - 284,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Remote AI Infra SRE — Reliability & Automation Lead
Remote AI Infra SRE — Reliability & Automation Lead

BridgeSource Utilities Solutions • United States

Hybrid
USD 140,000 - 190,000
Staff SRE: Reliability Architect for AI-Driven Platform
Staff SRE: Reliability Architect for AI-Driven Platform

Slope • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Benefits package