Site Reliability Engineer, AI Platform

United States Digital Space LLC

Paris (TX)

Hybrid

USD 80,000 - 111,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Algolia is seeking a Site Reliability Engineer to strengthen production fundamentals and automate complex systems at scale. You will own and operate Kubernetes-based platforms, improve reliability with observability, and drive deployment automation across AI-related workloads.

The role requires hands-on experience across cloud providers and strong scripting skills. Join a distributed team with offices in Paris, NYC, London, Sydney and Bucharest, and the option for remote or hybrid-remote work.

Qualifications

  • Solid hands-on Kubernetes knowledge, including workloads, resource management, and production operations.
  • Strong experience with Infrastructure as Code, and the lifecycle of cloud infrastructure.
  • Solid experience building and operating CI/CD pipelines and automated deployment workflows.
  • Hands-on experience with at least one major cloud provider: GCP, AWS or Azure.
  • Good understanding of networking, distributed systems and reliability engineering.
  • Experience with monitoring, observability and troubleshooting production systems.
  • Strong automation mindset and the ability to take ownership of well-defined production systems and progressively tackle more complex problems.
  • Excellent written and spoken English.

Responsibilities

  • Build and operate production infrastructure supporting AI-related workloads and services.
  • Operate and improve highly available Kubernetes-based platforms.
  • Improve reliability through SLOs, observability, alerting and capacity management.
  • Investigate production issues and turn findings into durable fixes and improvements.
  • Work across networking, databases, compute and service infrastructure.
  • Improve CI/CD pipelines, deployment automation and developer experience.
  • Build and maintain infrastructure using Infrastructure as Code.
  • Participate in on-call, incident response and operational improvements.
  • Collaborate with experienced engineers across AI Platform and progressively take ownership of broader production areas.

Skills

Kubernetes
Infrastructure as Code
CI/CD pipelines
GCP
AWS
Azure
Networking
Observability
Automation
English proficiency

Tools

Go
Python

Job description

the company is the retrieval intelligence layer that turns intent into trusted, decision-grade outcomes. Powering more than 1.7 trillion queries a year for over 18,000 customers with millisecond latency and 99.999% reliability, we are the recognized leader for Search and Product Discovery by top industry analyst firms. The the company platform turns a company's products, content, and business rules into data that humans, applications and AI agents can act upon. The result is trusted customer experiences with stronger conversions for measurable business impact.

the company was built to help users deliver intuitive search experiences across websites and mobile applications. Our Search API serves thousands of customers in more than 100 countries, answering billions of queries every month.

Join AI Platform: Powering AI in Production

AI Platform builds and operates the shared production foundations supporting the company's evolving AI ecosystem.

The team works at the intersection of Site Reliability Engineering, cloud infrastructure, software engineering and AI, helping engineering teams bring AI-powered capabilities to production reliably, securely and efficiently. Our scope includes Kubernetes, cloud infrastructure, CI/CD, networking, databases, observability, reliability, FinOps and production operations.

We are looking for a Site Reliability Engineer with strong production fundamentals who enjoys solving operational problems, automating repetitive work and progressively taking ownership of complex systems at scale.

YOU WILL:
  • Build and operate production infrastructure supporting AI-related workloads and services
  • Operate and improve highly available Kubernetes-based platforms
  • Improve reliability through SLOs, observability, alerting and capacity management
  • Investigate production issues and turn findings into durable fixes and improvements
  • Work across networking, databases, compute and service infrastructure
  • Improve CI/CD pipelines, deployment automation and developer experience
  • Build and maintain infrastructure using Infrastructure as Code
  • Participate in on-call, incident response and operational improvements
  • Collaborate with experienced engineers across AI Platform and progressively take ownership of broader production areas
YOU MIGHT BE A FIT IF YOU HAVE:
  • Solid hands-on Kubernetes knowledge, including workloads, resource management, and production operations
  • Strong experience with Infrastructure as Code, and the lifecycle of cloud infrastructure
  • Solid experience building and operating CI/CD pipelines and automated deployment workflows
  • Hands-on experience with at least one major cloud provider: GCP, AWS or Azure
  • Good understanding of networking, distributed systems and reliability engineering
  • Experience with monitoring, observability and troubleshooting production systems
  • Strong automation mindset and the ability to take ownership of well-defined production systems and progressively tackle more complex problems
  • Excellent written and spoken English
NICE TO HAVE:
  • Go and/or Python engineering experience
  • Exposure to AI/ML infrastructure and inferences
  • Comfortable working AI-first, using coding agents, agentic workflows and AI-assisted debugging to accelerate engineering and operations

the company does not discriminate on the basis of race, color, religion, sex, age, national origin, military status, veteran status, disability status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.

The annual base salary compensation range for this role reflects market pay data within this location. The exact compensation offered for this role may vary depending on specific location and job-related knowledge, technical skills, and experience; and is only one part of our Total Rewards philosophy to compensate and recognize employees for their work.

Base Salary Pay Range

€69.768—€96.900 EUR

FLEXIBLE WORKPLACE STRATEGY:

the company’s flexible workplace model is designed to empower all Algolians to fulfill our mission to power search and discovery with ease. We place an emphasis on an individual’s impact, contribution, and output, over their physical location. the company is a high-trust environment and many of our team members have the autonomy to choose where they want to work and when.

We have a global presence with offices in Paris, NYC, London, Sydney and Bucharest, however we also offer many of our team members the option to work remotely either as fully remote or hybrid-remote employees. Positions listed as \"Remote\" are only available for remote work within the specified country. Positions listed within a specific city are only available in that location - depending on the role it may be available with either a hybrid-remote or in-office schedule.WE’RE LOOKING FOR SOMEONE WHO CAN LIVE OUR VALUES:

  • GRIT - Problem-solving and perseverance capability in an ever-changing and growing environment.
  • TRUST - Willingness to trust our co-workers and to take ownership.
  • CANDOR - Ability to receive and give constructive feedback.
  • CARE - Genuine care about other team members, our clients and the decisions we make in the company.
  • HUMILITY - Aptitude for learning from others, putting ego aside.

We’re looking for talented, passionate people to help build the world’s best search and discovery technology. We value autonomy, diversity, and collaboration. We’re committed to creating an inclusive workplace where everyone is respected and supported—regardless of race, age, ancestry, religion, sex, gender identity, sexual orientation, marital status, color, veteran status, disability, or socioeconomic background.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, IaaS
Site Reliability Engineer, IaaS

United States Digital Space LLC • Paris (TX)

Hybrid
USD 140,000 - 190,000
Senior Site Reliability Engineer - Search
Senior Site Reliability Engineer - Search

United States Digital Space LLC • United States

Remote
USD 81,000 - 113,000
Site Reliability Engineer, IaaS
Site Reliability Engineer, IaaS

Algolia • United States

Hybrid
USD 65,000 - 90,000
L&D Program Manager New New York, New York
L&D Program Manager New New York, New York

Algolia, Inc. • New York (NY), Northern (KY)

On-site
USD 100,000 - 110,000
People Programs Manager
People Programs Manager

Algolia • New York (NY)

Hybrid
USD 100,000 - 110,000
Flexible workplace strategy
People Programs Manager New New York, New York
People Programs Manager New New York, New York

Algolia, Inc. • New York (NY), Northern (KY)

On-site
USD 100,000 - 110,000
Strategic Account Executive, Install Base
Strategic Account Executive, Install Base

Triwill Group • Northern (KY)

Hybrid
USD 296,000 - 320,000
Strategic Account Executive, Expansion & Renewals New Remote - United States
Strategic Account Executive, Expansion & Renewals New Remote - United States

Algolia, Inc. • Northern (KY)

Remote
USD 296,000 - 320,000
Flexible workplace
Employee Experience Specialist
Employee Experience Specialist

Triwill Group • New York (NY)

Hybrid
USD 100,000 - 110,000
Strategic Account Executive, Install Base
Strategic Account Executive, Install Base

Precision Labs • Northern (KY)

Hybrid
USD 296,000 - 320,000