Senior Site Reliability Engineer for Fuse team

bloomreach

United States

Hybrid

USD 81,000 - 127,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity-style awards
Education budget
Parental leave 26 weeks
Volunteer days
Well-being programs

Job summary

Bloomreach is seeking a senior Site Reliability Engineer to lead reliability and operability for a data platform team powering catalog and item-data services for personalization, search, and AI-enabled features. The role is based in Bratislava, Prague, or Brno, with hybrid options or remote work across Czechia or Slovakia.

You will own SLIs/SLOs, build end-to-end observability, maintain dashboards and alerts, and collaborate with backend, QA, product, and management to reduce noisy alerts and

Qualifications

  • Senior-level experience operating production distributed systems.
  • Hands-on with PostgreSQL, Bigtable, Elasticsearch, and Solr.
  • Experience defining SLIs, SLOs, and error budgets; turning ambiguities into measurable improvements.
  • Proficient with Grafana, Prometheus, OpenTelemetry, PagerDuty, and cloud platforms (GCP).
  • Ability to lead end-to-end reliability initiatives independently.

Responsibilities

  • Own and improve the reliability posture of services, workers, APIs, queues, storage, and pipelines.
  • Define SLIs, SLOs, and error budgets for APIs, asynchronous jobs, catalog data freshness, and indexing layers.
  • Build end-to-end observability so engineers can trace workspaces, item collections, catalogs, and jobs across services.
  • Maintain dashboards, alerts, and health views using Grafana, Prometheus-compatible metrics, OpenTelemetry, and cloud tooling.
  • Improve capacity planning and autoscaling through telemetry, queue depth, throughput, latency, and traffic patterns.
  • Partner with backend engineers, QA, product and management to reduce noisy alerts and tie signals to customer impact.

Skills

Senior engineer
Distributed systems
Observability
Leadership

Tools

PostgreSQL
Bigtable
Elasticsearch
Solr
Grafana
Prometheus
OpenTelemetry
PagerDuty
GCP

Job description

Role overview

A senior Site Reliability Engineer is needed to lead reliability and operability for a data platform team that owns catalog and item-data services feeding personalization, search, recommendations, and AI-agent products. The role centers on making distributed data pipelines observable, scalable, and safe to release across the full path from ingestion through storage and indexing to consumer products. The position is based in Bratislava, Prague, or Brno, with remote options in Czechia or Slovakia on a full-time basis.

Responsibilities
  • Own and improve the reliability posture of services, workers, APIs, queues, storage systems, and destination synchronization pipelines powering catalog and item data.
  • Define meaningful SLIs, SLOs, and error budgets for customer-facing APIs, asynchronous jobs, catalog data freshness, destination synchronization, and indexing layers.
  • Build end-to-end observability so engineers can trace workspaces, item collections, catalogs, and jobs across services without manual log correlation.
  • Maintain dashboards, alerts, and service-health views using Grafana, Prometheus-compatible metrics, OpenTelemetry, PagerDuty, and cloud-platform tooling.
  • Improve capacity planning and autoscaling through workload telemetry, queue depth, throughput, latency, memory, storage growth, and customer traffic patterns.
  • Partner with backend engineers, QA, product, and engineering management to reduce noisy alerts and replace symptom-based monitoring with signals tied to customer impact.
Requirements
  • Senior-level experience operating production distributed systems, including asynchronous job execution, queues, and large-scale storage and indexing components.
  • Hands-on expertise with PostgreSQL, Bigtable, Elasticsearch, and Solr in production environments.
  • Practical experience defining SLIs, SLOs, and error budgets and turning ambiguous operational problems into measurable improvements.
  • Strong skills with Grafana, Prometheus, OpenTelemetry, PagerDuty, and a major cloud platform (GCP experience preferred).
  • Ability to lead end-to-end reliability initiatives independently with minimal day-to-day guidance.
  • Comfort using AI-assisted coding tools thoughtfully while retaining ownership of correctness, security, and production outcomes.
Nice to have
  • Background in catalog or product-data systems, schema evolution, or backward-compatible API design.
  • Familiarity with destination synchronization, migrations, and multi-tenant data lifecycle management.
Benefits and work setup
  • Hybrid options from Central European offices (Bratislava, Prague, Brno) or remote within Czechia or Slovakia.
  • Support for professional development through an annual education budget, internal workshops, and access to a communication coach.
  • Well-being offerings including an employee assistance program, meditation app subscription, additional quarterly DisConnect days, and sports or yoga opportunities.
  • Extended parental leave up to 26 calendar weeks for primary caregivers, plus paid volunteering days.
  • Equity-style awards, employee referral bonuses, and a company performance bonus tied to organizational results.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Analyst
Senior Data Analyst

Hidden Jobs • United States

Hybrid
USD 90,000 - 140,000
Gen AI Academy access
Learning platform
International collaboration
+1
Cloud Engineer (AWS)
Cloud Engineer (AWS)

Digitalents Ukraine • Indiana (PA)

Hybrid
Five extra vacation days
Sick leave compensation 100%
Yearly bonus up to 10%
+8
Senior Staff Software Engineer, Data Delivery
Senior Staff Software Engineer, Data Delivery

Jobgether • United States

On-site
USD 240,000 - 330,000
Health benefits
401(k)/pension programs
Paid time off
+3
Middle/Senior Java Developer
Middle/Senior Java Developer

veeamsoftware • United States

Hybrid
USD 110,000 - 170,000
Hybrid work
Generous paid time off
Private medical insurance
+6
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Remote Jobs • United States

Remote
USD 150,000 - 210,000
Competitive Salary
Quarterly Bonuses
Unlimited PTO
+7
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

Juniper Square • United States

On-site
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+3
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Software Engineer, Backend - Data Layer
Software Engineer, Backend - Data Layer

Camunda • Raleigh (NC)

On-site
USD 110,000 - 160,000
Home office budget
Co-working space
Flexible time off
+11
Site Reliability Engineer
Site Reliability Engineer

Esker • Town of Middleton (WI)

On-site
USD 95,000 - 119,000
Student loan repayment assistance
Flexible work schedule
Summer hours
+5
Senior Fullstack Software Engineer (Backend-focused, Java/Spring Boot)
Senior Fullstack Software Engineer (Backend-focused, Java/Spring Boot)

FusionHit • United States

Remote
USD 120,000 - 180,000
Great Place to Work
Work from home
Flexible Schedule
+2