Remote SRE for AI Platform - Scale & Reliability

United States Digital Space LLC

Paris (TX)

On-site

USD 80,000 - 111,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Algolia is seeking a Site Reliability Engineer to strengthen production fundamentals and automate complex systems at scale. You will own and operate Kubernetes-based platforms, improve reliability with observability, and drive deployment automation across AI-related workloads.

The role requires hands-on experience across cloud providers and strong scripting skills. Join a distributed team with offices in Paris, NYC, London, Sydney and Bucharest, and the option for remote or hybrid-remote work.

Qualifications

  • Solid hands-on Kubernetes knowledge, including workloads, resource management, and production operations.
  • Strong experience with Infrastructure as Code, and the lifecycle of cloud infrastructure.
  • Solid experience building and operating CI/CD pipelines and automated deployment workflows.
  • Hands-on experience with at least one major cloud provider: GCP, AWS or Azure.
  • Good understanding of networking, distributed systems and reliability engineering.
  • Experience with monitoring, observability and troubleshooting production systems.
  • Strong automation mindset and the ability to take ownership of well-defined production systems and progressively tackle more complex problems.
  • Excellent written and spoken English.

Responsibilities

  • Build and operate production infrastructure supporting AI-related workloads and services.
  • Operate and improve highly available Kubernetes-based platforms.
  • Improve reliability through SLOs, observability, alerting and capacity management.
  • Investigate production issues and turn findings into durable fixes and improvements.
  • Work across networking, databases, compute and service infrastructure.
  • Improve CI/CD pipelines, deployment automation and developer experience.
  • Build and maintain infrastructure using Infrastructure as Code.
  • Participate in on-call, incident response and operational improvements.
  • Collaborate with experienced engineers across AI Platform and progressively take ownership of broader production areas.

Skills

Kubernetes
Infrastructure as Code
CI/CD pipelines
GCP
AWS
Azure
Networking
Observability
Automation
English proficiency

Tools

Go
Python

Job description

Algolia is seeking a Site Reliability Engineer to strengthen production fundamentals and automate complex systems at scale. You will own and operate Kubernetes-based platforms, improve reliability with observability, and drive deployment automation across AI-related workloads.

The role requires hands-on experience across cloud providers and strong scripting skills. Join a distributed team with offices in Paris, NYC, London, Sydney and Bucharest, and the option for remote or hybrid-remote work.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2
Site Reliability Engineer, Cloud & Kubernetes — Remote
Site Reliability Engineer, Cloud & Kubernetes — Remote

United States Digital Space LLC • Paris (TX)

On-site
USD 140,000 - 190,000
Senior SRE – AI Search, Remote-Optional
Senior SRE – AI Search, Remote-Optional

United States Digital Space LLC • United States

Remote
USD 81,000 - 113,000
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops

Arcoro Holdings Corp • Phoenix (AZ), Northern (KY)

Hybrid
USD 200,000 - 220,000
Remote Work
401(k) with Company match
Flexible PTO and Company-paid holidays
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
AI Platform SRE: Reliability, Observability & Scale
AI Platform SRE: Reliability, Observability & Scale

Schonfeld • New York (NY)

On-site
USD 175,000 - 225,000
Remote AI Infrastructure SRE — Kubernetes & Reliability
Remote AI Infrastructure SRE — Kubernetes & Reliability

Andromeda • San Francisco (CA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer, AI Platform
Site Reliability Engineer, AI Platform

United States Digital Space LLC • Paris (TX)

Hybrid
USD 80,000 - 111,000
Senior SRE: AI Cloud Reliability & Observability (Remote)
Senior SRE: AI Cloud Reliability & Observability (Remote)

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Remote Senior SRE: Build Reliable, Scalable AI Infra
Remote Senior SRE: Build Reliable, Scalable AI Infra

Runware • Town of Sweden (NY)

On-site
USD 140,000 - 190,000
Generous paid time off
Meaningful stock options
Remote-first setup
+3