Global Remote SRE for AI Infrastructure & Kubernetes

Andromeda Cluster

San Francisco (CA)

Hybrid

USD 120,000 - 160,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

A cutting-edge AI infrastructure company is seeking a Site Reliability Engineer to manage Kubernetes clusters and improve the reliability of critical systems. The ideal candidate will have 5+ years of experience in SRE or DevOps, strong Linux and Kubernetes expertise, and skills in automation and Infrastructure-as-Code. This role offers the opportunity to shape the future of scalable AI infrastructure, working closely with both customers and technical teams in a dynamic environment.

Qualifications

  • 5+ years experience in SRE, DevOps, or infrastructure engineering roles.
  • Strong Linux systems and networking fundamentals.
  • Deep experience with Kubernetes and container orchestration at scale.

Responsibilities

  • Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers.
  • Build automation and tooling to streamline cluster deployments and integrations.
  • Collaborate with engineering and product teams to plan and deliver infrastructure for new services.

Skills

SRE experience
Linux systems knowledge
Kubernetes expertise
Infrastructure-as-Code proficiency
Scripting skills in Python, Go, or Bash
Automation skills
Experience with observability tools

Tools

Terraform
Ansible
Prometheus
Grafana

Job description

A cutting-edge AI infrastructure company is seeking a Site Reliability Engineer to manage Kubernetes clusters and improve the reliability of critical systems. The ideal candidate will have 5+ years of experience in SRE or DevOps, strong Linux and Kubernetes expertise, and skills in automation and Infrastructure-as-Code. This role offers the opportunity to shape the future of scalable AI infrastructure, working closely with both customers and technical teams in a dynamic environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE — AI-Driven Cloud Reliability
Senior SRE — AI-Driven Cloud Reliability

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
SRE: AI/ML Infra on Kubernetes, AWS & Terraform
SRE: AI/ML Infra on Kubernetes, AWS & Terraform

Deepgram • United States

Hybrid
USD 120,000 - 150,000
Site Reliability Engineer - Kubernetes & Cloud
Site Reliability Engineer - Kubernetes & Cloud

Hydrolix • United States

On-site
USD 110,000 - 150,000
Senior Kubernetes SRE for AI Infrastructure (Bare-Metal)
Senior Kubernetes SRE for AI Infrastructure (Bare-Metal)

Moonlite • Chicago (IL)

Hybrid
USD 165,000 - 225,000
Competitive total compensation
401(k) match
Fully covered health insurance premiums
Senior Site Reliability Engineer: Cloud, Kubernetes & CI/CD
Senior Site Reliability Engineer: Cloud, Kubernetes & CI/CD

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Remote AI Infrastructure SRE — Kubernetes & Reliability
Remote AI Infrastructure SRE — Kubernetes & Reliability

Andromeda • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior SRE: Kubernetes, GPU Infra & ML Ops Leader
Senior SRE: Kubernetes, GPU Infra & ML Ops Leader

Gruve • Redwood City (CA)

On-site
USD 120,000 - 150,000
Senior SRE: AI-Powered Reliability on AWS & Kubernetes
Senior SRE: AI-Powered Reliability on AWS & Kubernetes

BetterUp • Austin (TX)

Hybrid
USD 147,000 - 185,000
Remote SRE & Platform Engineer - Hybrid Cloud & Kubernetes
Remote SRE & Platform Engineer - Hybrid Cloud & Kubernetes

Stash Talent Services • Virginia (MN)

Remote
USD 80,000 - 100,000
Remote AI Reliability Engineer (SRE) for Gen AI Systems
Remote AI Reliability Engineer (SRE) for Gen AI Systems

DeWinter Group • Campbell (CA)

Remote