Global Remote SRE for AI Infrastructure & Kubernetes
Andromeda Cluster
San Francisco (CA)
Hybrid
USD 120,000 - 160,000
Full time
14 days+
Application generator
Turn this role into an interview — a resume and cover letter built around what this employer wants.
Get past ATS filters
Job summary
A cutting-edge AI infrastructure company is seeking a Site Reliability Engineer to manage Kubernetes clusters and improve the reliability of critical systems. The ideal candidate will have 5+ years of experience in SRE or DevOps, strong Linux and Kubernetes expertise, and skills in automation and Infrastructure-as-Code. This role offers the opportunity to shape the future of scalable AI infrastructure, working closely with both customers and technical teams in a dynamic environment.
Qualifications
5+ years experience in SRE, DevOps, or infrastructure engineering roles.
Strong Linux systems and networking fundamentals.
Deep experience with Kubernetes and container orchestration at scale.
Responsibilities
Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers.
Build automation and tooling to streamline cluster deployments and integrations.
Collaborate with engineering and product teams to plan and deliver infrastructure for new services.
Skills
SRE experience
Linux systems knowledge
Kubernetes expertise
Infrastructure-as-Code proficiency
Scripting skills in Python, Go, or Bash
Automation skills
Experience with observability tools
Tools
Terraform
Ansible
Prometheus
Grafana
Job description
A cutting-edge AI infrastructure company is seeking a Site Reliability Engineer to manage Kubernetes clusters and improve the reliability of critical systems. The ideal candidate will have 5+ years of experience in SRE or DevOps, strong Linux and Kubernetes expertise, and skills in automation and Infrastructure-as-Code. This role offers the opportunity to shape the future of scalable AI infrastructure, working closely with both customers and technical teams in a dynamic environment.