Site Reliability Engineer, IaaS

Algolia

United States

Hybrid

USD 65,000 - 90,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Algolia is seeking a Site Reliability Engineer for the IaaS team to help build the next generation of production infrastructure. You will contribute to cloud foundations, Kubernetes, automation, reliability, and large-scale operations across cloud environments.

This is a hands-on role focused on building, operating, and improving scalable systems with emphasis on safety and observability. You will work with cross-functional teams to deliver secure, cost-aware platform capabilities and

Qualifications

  • Hands-on production knowledge of AWS or GCP.
  • Practical Kubernetes knowledge and an interest in operating it in production.
  • Familiarity with infrastructure as code, ideally Terraform.
  • Programming or scripting skills in Python, Go, or an equivalent language.
  • Strong Linux and networking fundamentals.
  • A strong interest in reliability, automation, and solving production problems.
  • Comfort adopting AI-assisted engineering tools, with sound judgement for critical production systems.
  • The ability to communicate clearly and work effectively with a distributed team.
  • Excellent spoken and written English skills.

Responsibilities

  • Build and improve Cloud Baseline capabilities, including identity and access, networking, security, resource inventory, tagging, and auditability.
  • Develop and maintain infrastructure as code and automation for cloud environments and Kubernetes infrastructure.
  • Contribute to reliable, repeatable cloud and cluster lifecycle operations.
  • Help build self-service capabilities, reusable modules, and clear documentation that make the safe path the easy path for platform consumers.
  • Reduce manual work and configuration drift through automation, testing, GitOps practices, and standardisation.
  • Use automation and AI-assisted engineering tools where appropriate to improve infrastructure analysis, documentation, and safe, repeatable changes.
  • Improve observability, monitoring, alerting, capacity management, and operational documentation.
  • Investigate production issues, participate in the on-call rotation, and turn lessons learned into lasting improvements.
  • Work with Infrastructure, Security, FinOps, and engineering teams to deliver reliable, secure, and cost-aware platform capabilities.

Skills

AWS/GCP
Kubernetes
Infrastructure as code
Terraform
Python/Go
Linux
Networking
Reliability
Communication
English

Tools

Argo CD
Helm
OPA
Kyverno

Job description

Algolia is the retrieval intelligence layer that turns intent into trusted, decision-grade outcomes. Powering more than 1.7 trillion queries a year for over 18,000 customers with millisecond latency and 99.999% reliability, we are the recognized leader for Search and Product Discovery by top industry analyst firms. The Algolia platform turns a company's products, content, and business rules into data that humans, applications and AI agents can act upon. The result is trusted customer experiences with stronger conversions for measurable business impact.

The team

The Infrastructure as a Service team is at the center of one of Algolia’s most consequential engineering transformations.

For years, Algolia has operated a production fleet of approximately 4,000 bare-metal servers to deliver the reliability, low latency, and scalability that our customers expect. We are now building the foundations of a unified cloud and Kubernetes platform designed to support Algolia’s growth for years to come.

. We are now building the foundations of a unified cloud and Kubernetes platform designed to support Algolia’s growth for years to come.

This is not a lift-and-shift project. It is an opportunity to rethink how Algolia provisions, secures, operates, observes, upgrades, and scales production infrastructure and to build it as a platform that engineers can safely consume, rather than a queue of manual requests.

The opportunity

As a Site Reliability Engineer in IaaS, you will help build the next generation of Algolia’s production infrastructure.

You will contribute to the cloud baseline and reliable lifecycle capabilities that enable teams to operate and migrate workloads safely on a cloud-native platform. You will work across cloud foundations, Kubernetes, automation, reliability, and large-scale production operations.

As a P3 engineer, you will be a hands-on contributor. You will build, operate, and improve production infrastructure while developing deep expertise in cloud, Kubernetes, reliability, and automation.

YOU WILL:
  • Build and improve Cloud Baseline capabilities, including identity and access, networking, security, resource inventory, tagging, and auditability.
  • Develop and maintain infrastructure as code and automation for cloud environments and Kubernetes infrastructure.
  • Contribute to reliable, repeatable cloud and cluster lifecycle operations.
  • Help build self-service capabilities, reusable modules, and clear documentation that make the safe path the easy path for platform consumers.
  • Reduce manual work and configuration drift through automation, testing, GitOps practices, and standardisation.
  • Use automation and AI-assisted engineering tools where appropriate to improve infrastructure analysis, documentation, and safe, repeatable changes.
  • Improve observability, monitoring, alerting, capacity management, and operational documentation.
  • Investigate production issues, participate in the on-call rotation, and turn lessons learned into lasting improvements.
  • Work with Infrastructure, Security, FinOps, and engineering teams to deliver reliable, secure, and cost-aware platform capabilities.
YOU MIGHT BE A FIT IF YOU HAVE:
  • Hands-on production knowledge of AWS or GCP.
  • Practical Kubernetes knowledge and an interest in operating it in production.
  • Familiarity with infrastructure as code, ideally Terraform.
  • Programming or scripting skills in Python, Go, or an equivalent language.
  • Strong Linux and networking fundamentals.
  • A strong interest in reliability, automation, and solving production problems.
  • Comfort adopting AI-assisted engineering tools, with sound judgement for critical production systems.
  • The ability to communicate clearly and work effectively with a distributed team.
  • Excellent spoken and written English skills.
NICE TO HAVE:
  • Familiarity with more than one public cloud provider.
  • Knowledge of GitOps or policy-as-code tooling, such as Argo CD, Helm, OPA, or Kyverno.
  • Experience with cloud migration, platform engineering, or large-scale infrastructure transformation.

Algolia does not discriminate on the basis of race, color, religion, sex, age, national origin, military status, veteran status, disability status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.

The annual base salary compensation range for this role reflects market pay data within this location. The exact compensation offered for this role may vary depending on specific location and job-related knowledge, technical skills, and experience; and is only one part of our Total Rewards philosophy to compensate and recognize employees for their work.

Base Salary Pay Range

€56.508 - €78.540 EUR

FLEXIBLE WORKPLACE STRATEGY:

Algolia’s flexible workplace model is designed to empower all Algolians to fulfill our mission to power search and discovery with ease. We place an emphasis on an individual’s impact, contribution, and output, over …

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, AI Platform
Site Reliability Engineer, AI Platform

United States Digital Space LLC • Paris (TX)

Hybrid
USD 80,000 - 111,000
Site Reliability Engineer, IaaS
Site Reliability Engineer, IaaS

United States Digital Space LLC • Paris (TX)

Hybrid
USD 140,000 - 190,000
Senior Software Engineer, Enterprise Commerce – Data (SFCC)
Senior Software Engineer, Enterprise Commerce – Data (SFCC)

Precision Labs • Northern (KY)

Hybrid
USD 163,000 - 214,000
People Programs Manager
People Programs Manager

Algolia • New York (NY)

Hybrid
USD 100,000 - 110,000
Flexible workplace strategy
L&D Program Manager New New York, New York
L&D Program Manager New New York, New York

Algolia, Inc. • New York (NY), Northern (KY)

On-site
USD 100,000 - 110,000
Employee Engagement Specialist
Employee Engagement Specialist

Algolia • New York (NY)

Hybrid
USD 100,000 - 110,000
Employee Experience Specialist
Employee Experience Specialist

Triwill Group • New York (NY)

Hybrid
USD 100,000 - 110,000
Site Reliability Engineer, Cloud & Kubernetes — Remote
Site Reliability Engineer, Cloud & Kubernetes — Remote

United States Digital Space LLC • Paris (TX)

On-site
USD 140,000 - 190,000
Employee Engagement Specialist
Employee Engagement Specialist

Visa Hunt • New York (NY)

Hybrid
USD 100,000 - 110,000
People Programs Manager New New York, New York
People Programs Manager New New York, New York

Algolia, Inc. • New York (NY), Northern (KY)

On-site
USD 100,000 - 110,000