Senior SRE, AI Platform: Scale Reliability & Kubernetes

United States Digital Space LLC

Paris (TX)

On-site

USD 80,000 - 111,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Algolia is seeking a Senior Site Reliability Engineer to own and evolve production systems, focusing on highly available Kubernetes-based platforms and AI-focused workloads. You will drive reliability, observability, and scalable infrastructure, collaborating across teams to deliver resilient services.

The role emphasizes independent problem ownership, strong automation, and mentorship, with a flexible workplace strategy that supports remote or hybrid work from multiple locations, including

Qualifications

  • At least one major cloud provider experience (GCP, AWS or Azure) is required.
  • Strong Kubernetes design and operations experience at scale.
  • Solid understanding of distributed systems and networking.
  • Experience operating business-critical systems with high availability and reliability.
  • Ability to own ambiguous, cross-team technical problems and drive them to outcomes.
  • Strong automation mindset balancing reliability, velocity and cost.
  • Excellent written and spoken English.

Responsibilities

  • Own and evolve production infrastructure supporting AI workloads and services at scale.
  • Design and operate Kubernetes-based platforms.
  • Drive reliability through SLOs, observability, capacity planning and production guardrails.
  • Lead complex production investigations and turn findings into durable architectural improvements.
  • Improve shared infrastructure across networking, databases, service communication and compute.
  • Build better CI/CD, progressive delivery, automation and developer experience.
  • Drive cloud infrastructure efficiency and FinOps initiatives.
  • Participate in and improve on-call and incident response.
  • Mentor engineers and raise the technical bar for reliability and production engineering.

Skills

Kubernetes
Cloud platforms
Distributed systems
Networking
Incident response
Automation
English fluency

Tools

Go
Python

Job description

Algolia is seeking a Senior Site Reliability Engineer to own and evolve production systems, focusing on highly available Kubernetes-based platforms and AI-focused workloads. You will drive reliability, observability, and scalable infrastructure, collaborating across teams to deliver resilient services.

The role emphasizes independent problem ownership, strong automation, and mentorship, with a flexible workplace strategy that supports remote or hybrid work from multiple locations, including

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote SRE for AI Platform - Scale & Reliability
Remote SRE for AI Platform - Scale & Reliability

United States Digital Space LLC • Paris (TX)

On-site
USD 80,000 - 111,000
Senior SRE: CI/CD, Kubernetes & Observability Leader
Senior SRE: CI/CD, Kubernetes & Observability Leader

United States Digital Space LLC • Paris (TX)

On-site
USD 80,000 - 111,000
Remote work options
Global offices
SRE, Cloud & Kubernetes Platform Engineer
SRE, Cloud & Kubernetes Platform Engineer

Algolia • United States

Hybrid
USD 65,000 - 90,000
Senior SRE, Cloud Baseline & Kubernetes Platform
Senior SRE, Cloud Baseline & Kubernetes Platform

United States Digital Space LLC • Paris (TX)

On-site
USD 140,000 - 210,000
Site Reliability Engineer, Cloud & Kubernetes — Remote
Site Reliability Engineer, Cloud & Kubernetes — Remote

United States Digital Space LLC • Paris (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior SRE: AI-Driven Platform Reliability & Scale
Senior SRE: AI-Driven Platform Reliability & Scale

Medallia • McLean (VA)

On-site
USD 129,000 - 190,000
Health benefits
401(k) matching
Paid parental leave
+1
Senior SRE – AI Search, Remote-Optional
Senior SRE – AI Search, Remote-Optional

United States Digital Space LLC • United States

Remote
USD 81,000 - 113,000
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule