Site Reliability Engineer, IaaS

United States Digital Space LLC

Paris (TX)

Hybrid

USD 140,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Algolia in the United States seeks a Site Reliability Engineer for the IaaS team to help build the next generation of production infrastructure. You will own cloud foundations, Kubernetes, and automation, delivering reliable, scalable platform capabilities for global customers.

The role requires hands-on AWS or GCP experience, Kubernetes in production, and IaC with Terraform, plus Python/Go scripting. A flexible remote/hybrid work model and strong collaboration across teams are offered.

Qualifications

  • Hands-on production knowledge of AWS or GCP.
  • Practical Kubernetes knowledge and production operations.
  • Familiarity with infrastructure as code (Terraform).
  • Programming or scripting in Python or Go.
  • Strong Linux and networking fundamentals.

Responsibilities

  • Build and improve Cloud Baseline capabilities, including identity and access, networking, security, resource inventory, tagging, and auditability.
  • Develop and maintain infrastructure as code and automation for cloud environments and Kubernetes infrastructure.
  • Contribute to reliable, repeatable cloud and cluster lifecycle operations.
  • Help build self-service capabilities, reusable modules, and clear documentation that make the safe path the easy path for platform consumers.
  • Reduce manual work and configuration drift through automation, testing, GitOps practices, and standardisation.
  • Use automation and AI-assisted engineering tools where appropriate to improve infrastructure analysis, documentation, and safe, repeatable changes.
  • Improve observability, monitoring, alerting, capacity management, and operational documentation.
  • Investigate production issues, participate in the on-call rotation, and turn lessons learned into lasting improvements.
  • Work with Infrastructure, Security, FinOps, and engineering teams to deliver reliable, secure, and cost-aware platform capabilities.

Tools

Terraform
Argo CD
Helm
OPA
Kyverno
AWS or GCP
Kubernetes
Python
Go
Linux fundamentals
Networking fundamentals
Automation
Reliability
AI tooling

Job description

The team

At the company, we’re proud to be a pioneer and market leader in AI Search, empowering 18,000+ businesses to deliver blazing-fast, predictive search and browse experiences at internet scale. Every week, we power over 30 billion search requests — four times more than Microsoft Bing, Yahoo, Baidu, Yandex, and DuckDuckGo combined.

In 2021, we raised $150 million in Series D funding, quadrupling our valuation to $2.25 billion. This strong foundation enables us to keep investing in our market-leading platform and serving incredible customers like Under Armour, PetSmart, Stripe, Gymshark, and Walgreens.

The team

The Infrastructure as a Service team is at the center of one of the company’s most consequential engineering transformations.

For years, the company has operated a production fleet of approximately 4,000 bare-metal servers to deliver the reliability, low latency, and scalability that our customers expect. We are now building the foundations of a unified cloud and Kubernetes platform designed to support the company’s growth for years to come.

. We are now building the foundations of a unified cloud and Kubernetes platform designed to support the company’s growth for years to come.

This is not a lift-and-shift project. It is an opportunity to rethink how the company provisions, secures, operates, observes, upgrades, and scales production infrastructure and to build it as a platform that engineers can safely consume, rather than a queue of manual requests.

The opportunity

As a Site Reliability Engineer in IaaS, you will help build the next generation of the company’s production infrastructure.

You will contribute to the cloud baseline and reliable lifecycle capabilities that enable teams to operate and migrate workloads safely on a cloud-native platform. You will work across cloud foundations, Kubernetes, automation, reliability, and large-scale production operations.

As a P3 engineer, you will be a hands-on contributor. You will build, operate, and improve production infrastructure while developing deep expertise in cloud, Kubernetes, reliability, and automation.

YOU WILL:
  • Build and improve Cloud Baseline capabilities, including identity and access, networking, security, resource inventory, tagging, and auditability.
  • Develop and maintain infrastructure as code and automation for cloud environments and Kubernetes infrastructure.
  • Contribute to reliable, repeatable cloud and cluster lifecycle operations.
  • Help build self-service capabilities, reusable modules, and clear documentation that make the safe path the easy path for platform consumers.
  • Reduce manual work and configuration drift through automation, testing, GitOps practices, and standardisation.
  • Use automation and AI-assisted engineering tools where appropriate to improve infrastructure analysis, documentation, and safe, repeatable changes.
  • Improve observability, monitoring, alerting, capacity management, and operational documentation.
  • Investigate production issues, participate in the on-call rotation, and turn lessons learned into lasting improvements.
  • Work with Infrastructure, Security, FinOps, and engineering teams to deliver reliable, secure, and cost-aware platform capabilities.
YOU MIGHT BE A FIT IF YOU HAVE:
  • Hands-on production knowledge of AWS or GCP.
  • Practical Kubernetes knowledge and an interest in operating it in production.
  • Familiarity with infrastructure as code, ideally Terraform.
  • Programming or scripting skills in Python, Go, or an equivalent language.
  • Strong Linux and networking fundamentals.
  • A strong interest in reliability, automation, and solving production problems.
  • Comfort adopting AI-assisted engineering tools, with sound judgement for critical production systems.
  • The ability to communicate clearly and work effectively with a distributed team.
  • Excellent spoken and written English skills.
NICE TO HAVE:
  • Familiarity with more than one public cloud provider.
  • Knowledge of GitOps or policy-as-code tooling, such as Argo CD, Helm, OPA, or Kyverno.
  • Experience with cloud migration, platform engineering, or large-scale infrastructure transformation.
FLEXIBLE WORKPLACE STRATEGY:

the company’s flexible workplace model is designed to empower all Algolians to fulfill our mission to power search and discovery with ease. We place an emphasis on an individual’s impact, contribution, and output, over their physical location. the company is a high-trust environment and many of our team members have the autonomy to choose where they want to work and when.

We have a global presence with offices in Paris, NYC, London, Sydney and Bucharest, however we also offer many of our team members the option to work remotely either as fully remote or hybrid-remote employees.

Positions listed as "Remote" are only available for remote work within the specified country. Positions listed within a specific city are only available in that location - depending on the role it may be available with either a hybrid-remote or in-office schedule.WE’RE LOOKING FOR SOMEONE WHO CAN LIVE OUR VALUES:
  • GRIT - Problem-solving and perseverance capability in an ever-changing and growing environment.
  • TRUST - Willingness to trust our co-workers and to take ownership.
  • CANDOR - Ability to receive and give constructive feedback.
  • CARE - Genuine care about other team members, our clients and the decisions we make in the company.
  • HUMILITY - Aptitude for learning from others, putting ego aside.

We’re looking for talented, passionate people to help build the world’s best search and discovery technology. We value autonomy, diversity, and collaboration. We’re committed to creating an inclusive workplace where everyone is respected and supported—regardless of race, age, ancestry, religion, sex, gender identity, sexual orientation, marital status, color, veteran status, disability, or socioeconomic background.

IMPORTANT NOTICE FOR CANDIDATES - Recruitment Fraud Notice

We’ve recently seen an increase in recruitment scams targeting job seekers. To help protect yourself, please keep the following in mind:

  • Our open positions may appear on third-party job boards, but the best way to apply safely is directly through our careers page.
  • All genuine communication from the company will come from an ----- email address. If you receive an email from someone claiming to work at the company who does not have an ----- email address, please do not respond or share any personal information.
  • We’ll never ask for payments, purchases, or financial details during the hiring process.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, AI Platform
Site Reliability Engineer, AI Platform

United States Digital Space LLC • Paris (TX)

Hybrid
USD 80,000 - 111,000
Senior Site Reliability Engineer - Search
Senior Site Reliability Engineer - Search

United States Digital Space LLC • United States

Remote
USD 81,000 - 113,000
Remote Solutions Engineer - AI Search Pre-Sales
Remote Solutions Engineer - AI Search Pre-Sales

Idlen Inc. • United States

Remote
Strategic Account Executive, Install Base New Remote - United States
Strategic Account Executive, Install Base New Remote - United States

Algolia, Inc. • Northern (KY)

Remote
USD 296,000 - 320,000
Employee Engagement Specialist
Employee Engagement Specialist

Algolia • New York (NY)

Hybrid
USD 100,000 - 110,000
Senior Software Engineer, Enterprise Commerce – Data (SFCC)
Senior Software Engineer, Enterprise Commerce – Data (SFCC)

United States Digital Space LLC • United States

Hybrid
USD 163,000 - 214,000
Solutions Engineer
Solutions Engineer

Idlen Inc. • United States

On-site
Employee Engagement Specialist
Employee Engagement Specialist

Visa Hunt • New York (NY)

Hybrid
USD 100,000 - 110,000
L&D Program Manager New New York, New York
L&D Program Manager New New York, New York

Algolia, Inc. • New York (NY), Northern (KY)

On-site
USD 100,000 - 110,000
Senior Software Engineer, Enterprise Commerce – Data (SFCC)
Senior Software Engineer, Enterprise Commerce – Data (SFCC)

Precision Labs • Northern (KY)

Hybrid
USD 163,000 - 214,000