Site Reliability Engineer - Datastores

Coralogix

United States

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Coralogix is seeking a Site Reliability Engineer to join our team and help build the next generation of our stream-based observability platform. You will lead critical infrastructure projects, optimize Kubernetes deployments, and champion reliability, performance, and cost-efficiency at scale.

Ideal candidates have 5+ years in DevOps/SRE, strong experience with Kubernetes, cloud platforms (AWS, GCP), and distributed datastores. You will mentor teams and drive automation and innovative solutions.

Qualifications

  • 5+ years in DevOps, SRE, platform engineering, or infrastructure roles.
  • Experience maintaining datastores in high-scale environments (Mongo, Cassandra, ClickHouse, Redis, OpenSearch).
  • Experience running stateful workloads in Kubernetes.
  • Proven experience managing large-scale cloud infrastructure (AWS, GCP).
  • Experience in incident response and troubleshooting distributed systems.
  • Some software engineering experience, preferably in Golang.
  • Passion for automation, performance tuning, and operational excellence.

Responsibilities

  • Act as a hands-on technical leader with deep expertise in relational DBs or other distributed datastores.
  • Serve as a go-to person in the team — leading through influence, not hierarchy.

Skills

DevOps
SRE
Platform engineering
Automation
Golang

Tools

Kubernetes
AWS
GCP
Kafka
Istio
MongoDB
Cassandra
ClickHouse
Redis
OpenSearch

Job description

Coralogix is a modern, full-stack observability platform transforming how businesses process and understand their data. Our unique architecture powers in-stream analytics without reliance on expensive indexing or hot storage. We specialize in comprehensive monitoring of logs, metrics, trace, and security events with features such as APM, RUM, SIEM, Kubernetes monitoring, and more, all enhancing operational efficiency and reducing observability spending by up to 70%.

Coralogix is seeking a Site Reliability Engineer to join our team and help build the next generation of our stream-based observability platform. We deliver real-time data analytics at scale to some of the world's leading tech companies.

The SRE team is responsible for the foundational infrastructure that powers Coralogix:

  • Kubernetes Infrastructure: Managing over 10,000 nodes across multiple cloud providers and regions. Coralogix production is 100% Kubernetes based
  • Support service owners in running over 1,000 instances of multiple datastore types, cloud-based and self-hosted (on Kubernetes).
  • Maintaining critical, large-scale clusters processing billions of events per second.
  • Automation & Operators: Building and maintaining both open-source and custom Kubernetes operators to manage complex stateful workloads like Kafka. Our tech stack is constantly evolving. It includes: Kubernetes, Go (Golang), AWS, GCP, Kafka, Istio, and more.
  • Write and maintain Kubernetes Controllers using frameworks like controller-runtime and KubeBuilder
Responsibilities
  • Act as a hands-on technical leader with deep expertise in relational DBs or other distributed datastores
  • Serve as a go-to person in the team --- leading through influence, not hierarchy.
  • Collaborate cross-functionally to refine requirements and propose innovative, scalable solutions.
  • Drive long-term, high-impact infrastructure projects across multiple teams, from design to implementation, within defined timelines.
  • Contribute to improving system reliability, performance, and cost-efficiency at scale.
Requirements
  • 5+ years of experience in DevOps, SRE, platform engineering, or infrastructure roles.
  • Experience with maintaining datastores in high-scale environments, whether relational DBs or other distributed datastores (Mongo, Cassandra, ClickHouse, Redis, OpenSearch).
  • Experience with running stateful workloads in Kubernetes.
  • Proven experience managing large-scale cloud infrastructure (AWS, GCP, etc.).
  • Experience in incident response and troubleshooting complex distributed systems.
  • Some software engineering experience, preferably in Golang.
  • Passion for automation, performance tuning, and operational excellence.
Cultural Fit

We're seeking candidates who are hungry, humble, and smart. Coralogix fosters a culture of innovation and continuous learning, where team members are encouraged to challenge the status quo and contribute to our shared mission. If you thrive in dynamic environments and are eager to shape the future of observability solutions.

Coralogix is an equal opportunity employer and encourages applicants from all backgrounds to apply.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (FedRAMP / Security)
Site Reliability Engineer (FedRAMP / Security)

Coralogix • New York (NY)

On-site
USD 150,000 - 190,000
Equal opportunity employer
Senior SRE: Datastores & Scale in Kubernetes
Senior SRE: Datastores & Scale in Kubernetes

Coralogix • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (FedRAMP / Security) – CA
Site Reliability Engineer (FedRAMP / Security) – CA

Coralogix • Los Angeles (CA)

On-site
USD 170,000 - 220,000
Healthcare benefits
Dental benefits
401(k) plan
+1
Site Reliability Engineer (FedRAMP / Security)
Site Reliability Engineer (FedRAMP / Security)

Coralogix, inc. • Boston (MA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer (FedRAMP / Security)
Site Reliability Engineer (FedRAMP / Security)

Coralogix, inc. • New York (NY)

On-site
USD 120,000 - 150,000
Site Reliability Engineer (FedRAMP / Security) – CA
Site Reliability Engineer (FedRAMP / Security) – CA

Coralogix, inc. • Los Angeles (CA)

On-site
USD 170,000 - 220,000
Healthcare benefits
401(k) plan and match
Paid time off
Site Reliability Engineer (FedRAMP / Security) – NY
Site Reliability Engineer (FedRAMP / Security) – NY

Coralogix, inc. • New York (NY)

On-site
USD 170,000 - 220,000
Comprehensive healthcare benefits
401(k) plan and match
Paid time off
Site Reliability Engineer
Site Reliability Engineer

Hydrolix • United States

On-site
USD 110,000 - 150,000
Solution Engineer – San Francisco
Solution Engineer – San Francisco

Coralogix • San Francisco (CA)

On-site
USD 220,000 - 260,000
Healthcare
Dental
Mental health benefits
+3
Senior Solutions Engineer — Pre-Sales (Observability)
Senior Solutions Engineer — Pre-Sales (Observability)

Coralogix • Dallas (TX)

On-site
USD 220,000 - 250,000