Senior Elasticsearch Engineer

your Jared

Northern (KY)

Hybrid

USD 150,000 - 230,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully remote

Job summary

Chess.com, the world’s largest chess platform, seeks a Senior Elasticsearch Engineer to own the full lifecycle of our search and analytics data platform. You’ll drive capacity planning, cluster architecture, performance tuning, and incident response for massive-scale Elasticsearch/OpenSearch clusters on bare-metal Kubernetes.

You’ll work hands-on with ILM/ISM policies, GitOps pipelines, and real-time decisions on replica allocation during outages.

Qualifications

  • 7+ years operating Elasticsearch at scale (multi-TB clusters, dozens of nodes, high write throughput).
  • Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state management.
  • Production experience with Kubernetes-based deployments (StatefulSets, storage, resource management).
  • Experience managing Elasticsearch and OpenSearch in production with informed trade-offs.

Responsibilities

  • Shard allocation strategy for write-heavy data streams at high throughput (millions of documents per minute).
  • Disk watermark management, retention policy tuning, and rollover orchestration for high-volume indices.
  • Performance optimization and I/O tuning on bare-metal nodes.
  • Write queue analysis, thread pool diagnostics, and shard rebalancing under load.
  • Capacity planning and growth forecasting across clusters.
  • On-call ownership for Elasticsearch-related incidents: cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades.
  • Real-time cluster triage and cross-team coordination during production incidents.
  • Post-mortem authoring and systemic reliability improvements.
  • Snapshot and disaster recovery management across clusters.
  • Elasticsearch-to-OpenSearch migration analysis and execution, including compatibility evaluation across ILM/ISM, security models, and plugin ecosystems.
  • Version upgrade planning and rolling restart orchestration with zero-downtime requirements.
  • End-to-end new cluster provisioning and onboarding.
  • Advise engineering teams on index design, mapping strategy, retention policies, and query optimization.
  • Manage Kibana and OpenSearch Dashboards access and configuration for internal consumers.
  • Define and maintain workload priority tiers across clusters.

Skills

Elasticsearch
OpenSearch
Kubernetes
GitOps
Java internals
Linux admin
Incident response
Stakeholder communication

Tools

ECK
ArgoCD
Flux
Helm
S3/GCS snapshots

Job description

About Us

Chess.com is one of the largest gaming sites in the world and the #1 platform for playing, learning, and enjoying chess.

We are a team of 600+ fully remote people in 60+ countries working hard to serve the global chess community. We are here to support 250M+ chess players worldwide with the best possible product, content, and tools to serve the community!

We are a tech company. A gaming company. A content company. And we do it all with passion and commitment to the game. Above all we prize our mission-driven, flat, life-celebrating, no-corporate culture, and we look forward to meeting you and learning more about what you can bring to the team.

About You

Chess.com is the world's largest chess platform with 235M+ members and ~20 million daily games. Our Elasticsearch and OpenSearch infrastructure underpins search, user activity, analytics, logging, and operational intelligence at massive scale, hundreds of terabytes across a dozen production clusters running on bare-metal Kubernetes.

We're looking for a Senior Elasticsearch Engineer who can own the full lifecycle of our search and analytics data platform: capacity planning, cluster architecture, performance tuning, incident response, migration strategy, and operational excellence. You'll be the single point of deep expertise across all Elasticsearch and OpenSearch clusters at Chess.com.

This is not a monitoring-from-dashboards role. You'll be hands-on with cluster internals, write ILM/ISM policies, push infrastructure changes through GitOps, and make real-time decisions about replica allocation when a cluster goes red.

What you'll do
  • Shard allocation strategy for write-heavy data streams at high throughput (millions of documents per minute)
  • Disk watermark management, retention policy tuning, and rollover orchestration for high-volume indices
  • Performance optimization and I/O tuning on bare-metal nodes
  • Write queue analysis, thread pool diagnostics, and shard rebalancing under load
  • Capacity planning and growth forecasting across clusters
  • On-call ownership for Elasticsearch-related incidents: cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades
  • Real-time cluster triage and cross-team coordination during production incidents
  • Post-mortem authoring and systemic reliability improvements
  • Snapshot and disaster recovery management across clusters
  • Elasticsearch-to-OpenSearch migration analysis and execution, including compatibility evaluation across ILM/ISM, security models, and plugin ecosystems
  • Version upgrade planning and rolling restart orchestration with zero-downtime requirements
  • End-to-end new cluster provisioning and onboarding
  • Advise engineering teams on index design, mapping strategy, retention policies, and query optimization
  • Manage Kibana and OpenSearch Dashboards access and configuration for internal consumers
  • Define and maintain workload priority tiers across clusters
Preferred Skills
  • 7+ years operating Elasticsearch at scale (multi-TB clusters, dozens of nodes, high write throughput)
  • Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state management
  • Production experience with ECK (Elastic Cloud on Kubernetes) or equivalent operator-based deployments
  • Proficiency with Kubernetes operations for stateful workloads (StatefulSets, persistent storage, resource management)
  • Hands-on Linux systems administration with a focus on storage and I/O performance
  • Experience managing both Elasticsearch and OpenSearch in production, including an informed opinion on their respective trade-offs
  • Incident command experience: ability to diagnose and mitigate cluster emergencies under pressure while communicating clearly to stakeholders
  • Git-based infrastructure management (GitOps): Helm charts, ArgoCD/Flux, infrastructure-as-code for cluster configuration
  • Fluency with the Elastic stack APIs: cluster administration, index templates, data streams, ILM policies, snapshot/restore
  • OpenSearch ISM policies and security plugin (fine-grained access control)
  • GCS or S3 snapshot repository configuration and cross-cluster replication
  • Grafana + Prometheus monitoring for Elasticsearch metrics
  • Kibana Discover, Dev Tools, and data view management at scale
  • Java internals relevant to Elasticsearch JVM tuning (heap sizing, GC tuning, circuit breakers)
  • Vault integration for secrets management in Kubernetes-deployed search clusters
  • Fluentd/Fluent Bit log pipeline configuration feeding OpenSearch
  • Hardware selection experience for search-optimized server configurations
  • Python or scripting for operational analysis and automation
  • Full autonomy. You are the Elasticsearch authority. You make the architecture calls, set the priorities, and own the outcomes.
  • Real scale. Hundreds of terabytes of data, billions of documents, millions of daily queries. The problems here don't exist at smaller companies.
  • Bare metal. No managed Elastic Cloud. You're operating directly on the hardware. This is hands-on engineering.
  • Strategic impact. Your decisions on ES vs. OpenSearch migration, cluster topology, and capacity planning directly affect product capabilities and infrastructure costs.
  • Small team, big trust. Chess.com runs lean. You won't be buried in process or approvals. Ship changes, fix problems, improve systems.
About the Opportunity
  • This is a full-time opportunity
  • We are 100% remote (work from anywhere!)

You can learn more about us here:

Engineering

Remote

Share on:

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Elasticsearch & OpenSearch Engineer (Remote)
Senior Elasticsearch & OpenSearch Engineer (Remote)

Chess.com • Orem (UT)

On-site
USD 150,000 - 210,000
Senior Elasticsearch Architect — Remote & High-Scale Ops
Senior Elasticsearch Architect — Remote & High-Scale Ops

your Jared • Northern (KY)

Hybrid
USD 150,000 - 230,000
Fully remote
Senior Software Engineer - Query Engine, Database Internals - Elasticsearch Elastic
Senior Software Engineer - Query Engine, Database Internals - Elasticsearch Elastic

Neura Market • Northern (KY)

Hybrid
USD 133,000 - 211,000
Stock program
401k with company match up to 6%
Flexible locations and schedules
+4
Senior Software Developer - Query Engine, Database Internals - Elasticsearch Elastic
Senior Software Developer - Query Engine, Database Internals - Elasticsearch Elastic

Neura Market • Mountain View (CA)

Hybrid
USD 93,000 - 147,000
Stock option program
RRSP matching (6%)
Flexible work locations
Senior Software Developer - Storage Engine - Elasticsearch
Senior Software Developer - Storage Engine - Elasticsearch

Elastic • United States

Remote
CAD 128,000 - 203,000
Health coverage
Flexible locations
Vacation days
+4
Senior Java Engineer - Core/Infra - Elasticsearch
Senior Java Engineer - Core/Infra - Elasticsearch

Referral Board • United States

On-site
USD 133,100 - 210,600
Health coverage
Flexible locations and schedules
Generous vacation days
+3
Senior Software Engineer - Query Engine, Database Internals - Elasticsearch
Senior Software Engineer - Query Engine, Database Internals - Elasticsearch

Elasticsearch B.V. • Northern (KY)

Hybrid
USD 133,000 - 211,000
Health coverage
Flexible locations
Vacation days
+3
Principal Software Engineer - Query Engine, Database Internals - Elasticsearch Elastic
Principal Software Engineer - Query Engine, Database Internals - Elasticsearch Elastic

Neura Market • Northern (KY)

Hybrid
USD 160,000 - 253,000
Senior Java Engineer - Distributed Systems, Serverless - Elasticsearch
Senior Java Engineer - Distributed Systems, Serverless - Elasticsearch

Elastic • Town of Greece (NY)

On-site
USD 69,000 - 110,000
Health coverage
Flexible schedules
Vacation days
+3
Principal Java Engineer - Distributed Systems, Serverless - Elasticsearch
Principal Java Engineer - Distributed Systems, Serverless - Elasticsearch

Referral Board • United States

On-site
USD 133,100 - 210,600
Stock program
401k with match
Generous vacation days
+3