Site Reliability Engineer, AI Platform

Algolia, Inc.

Paris

Hybrid

EUR 80,000 - 120,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Algolia, Inc. is seeking a Site Reliability Engineer to join the AI Platform. You will deploy and run robust production infrastructure for AI workloads and maintain a Kubernetes-based platform at scale.

You will improve reliability with SLOs, observability, and automation while collaborating with AI Platform engineers to own broader production areas over time.

Qualifications

  • Solid hands-on Kubernetes knowledge, including workloads, resource management, and production operations.
  • Strong experience with Infrastructure as Code, and the lifecycle of cloud infrastructure.
  • Solid experience building and operating CI/CD pipelines and automated deployment workflows.
  • Hands-on experience with at least one major cloud provider: GCP, AWS or Azure.
  • Good understanding of networking, distributed systems and reliability engineering.
  • Experience with monitoring, observability and troubleshooting.

Responsibilities

  • Build and operate production infrastructure supporting AI-related workloads and services.
  • Operate and improve highly available Kubernetes-based platforms.
  • Improve reliability through SLOs, observability, alerting and capacity management.
  • Investigate production issues and turn findings into durable fixes and improvements.
  • Work across networking, databases, compute and service infrastructure.
  • Improve CI/CD pipelines, deployment automation and developer experience.
  • Build and maintain infrastructure using Infrastructure as Code.
  • Participate in on-call, incident response and operational improvements.
  • Collaborate with experienced engineers across AI Platform and progressively take ownership of broader production areas.

Skills

Kubernetes
Infrastructure as Code
CI/CD
Cloud platforms
Networking
Observability
Troubleshooting

Tools

Kubernetes

Job description

Algolia is the retrieval intelligence layer that turns intent into trusted, decision-grade outcomes. Powering more than 1.7 trillion queries a year for over 18,000 customers with millisecond latency and 99.999% reliability, we are the recognized leader for Search and Product Discovery by top industry analyst firms. The Algolia platform turns a company's products, content, and business rules into data that humans, applications and AI agents can act upon. The result is trusted customer experiences with stronger conversions for measurable business impact.

Algolia was built to help users deliver intuitive search experiences across websites and mobile applications. Our Search API serves thousands of customers in more than 100 countries, answering billions of queries every month.

Join AI Platform: Powering AI in Production

AI Platform builds and operates the shared production foundations supporting Algolia's evolving AI ecosystem.

The team works at the intersection of Site Reliability Engineering, cloud infrastructure, software engineering and AI, helping engineering teams bring AI-powered capabilities to production reliably, securely and efficiently. Our scope includes Kubernetes, cloud infrastructure, CI/CD, networking, databases, observability, reliability, FinOps and production operations.

We are looking for a Site Reliability Engineer with strong production fundamentals who enjoys solving operational problems, automating repetitive work and progressively taking ownership of complex systems at scale.

YOU WILL:
  • Build and operate production infrastructure supporting AI-related workloads and services
  • Operate and improve highly available Kubernetes-based platforms
  • Improve reliability through SLOs, observability, alerting and capacity management
  • Investigate production issues and turn findings into durable fixes and improvements
  • Work across networking, databases, compute and service infrastructure
  • Improve CI/CD pipelines, deployment automation and developer experience
  • Build and maintain infrastructure using Infrastructure as Code
  • Participate in on-call, incident response and operational improvements
  • Collaborate with experienced engineers across AI Platform and progressively take ownership of broader production areas
YOU MIGHT BE A FIT IF YOU HAVE:
  • Solid hands-on Kubernetes knowledge, including workloads, resource management, and production operations
  • Strong experience with Infrastructure as Code, and the lifecycle of cloud infrastructure
  • Solid experience building and operating CI/CD pipelines and automated deployment workflows
  • Hands-on experience with at least one major cloud provider: GCP, AWS or Azure
  • Good understanding of networking, distributed systems and reliability engineering
  • Experience with monitoring, observability and troubleshooting
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, AI Platform New Paris, France
Site Reliability Engineer, AI Platform New Paris, France

Algolia, Inc. • Paris

On-site
EUR 70,000 - 95,000
Flexible workplace
Remote or hybrid options
Site Reliability Engineer, AI Platform
Site Reliability Engineer, AI Platform

Engg • Paris

On-site
EUR 70,000 - 97,000
Senior Site Reliability Engineer, AI Platform
Senior Site Reliability Engineer, AI Platform

Algolia, Inc. • Paris

Hybrid
EUR 70,000 - 97,000
Senior Site Reliability Engineer, AI Platform
Senior Site Reliability Engineer, AI Platform

Algolia • Paris

Hybrid
EUR 70,000 - 97,000
Senior Site Reliability Engineer, AI Platform New Paris, France
Senior Site Reliability Engineer, AI Platform New Paris, France

Algolia, Inc. • Paris

On-site
EUR 90,000 - 140,000
Senior Site Reliability Engineer, IaaS
Senior Site Reliability Engineer, IaaS

Algolia, Inc. • Paris

Remote
EUR 110,000 - 165,000
Flexible workplace model
Remote or hybrid options
Site Reliability Engineer, IaaS
Site Reliability Engineer, IaaS

Engg • Paris

On-site
EUR 75,000 - 110,000
Remote work options
Global offices
Senior SRE: AI Platform, Kubernetes & Cloud Reliability
Senior SRE: AI Platform, Kubernetes & Cloud Reliability

Algolia, Inc. • Paris

Hybrid
EUR 70,000 - 97,000
AI Platform SRE: Scale Kubernetes & Automation
AI Platform SRE: Scale Kubernetes & Automation

Algolia, Inc. • Paris

Hybrid
EUR 70,000 - 95,000
Flexible workplace
Remote or hybrid options
Senior SRE, AI Platform — Production Reliability (Remote)
Senior SRE, AI Platform — Production Reliability (Remote)

Algolia • Paris

Hybrid
EUR 70,000 - 97,000