Remote Site Reliability Engineer — Global Scale

chess

United States

Remote

USD 150,000 - 190,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Chess.com is seeking a Site Reliability Engineer to ensure stability, performance, and scalability of our global gaming platform infrastructure. You will support millions of concurrent users, own on-call rotations, and drive incident response to maintain high availability.

You’ll collaborate across engineering teams to implement infrastructure-as-code, lead hybrid cloud migrations, and optimize monitoring, security, and capacity planning for a robust, scalable system.

Qualifications

  • Degree in computer science, engineering or related field.
  • Minimum 5 years in site reliability, DevOps or infra roles.
  • Strong UNIX/Linux experience and CLI proficiency.
  • Experience with cloud platforms (GCP/AWS/Azure) and IaC tools.
  • Hands-on with configuration management (Ansible, etc.).
  • Knowledge of networking, DNS, TCP/IP, HTTP/HTTPS.
  • Experience with Docker and Kubernetes in production.
  • Proficient in observability tooling and monitoring.

Responsibilities

  • Design multi-regional, resilient infra for millions of sessions daily.
  • Lead hybrid cloud migration, blending bare‑metal and cloud resources.
  • Own on-call rotation and incident response with high availability SLAs.
  • Architect monitoring/alerting to detect bottlenecks early.
  • Collaborate on IaC and deployment pipelines for CI/CD.
  • Optimize capacity planning and resource use across regions.
  • Establish security protocols and data protection measures.
  • Partner with engineers on scalable real-time processing solutions.
  • Drive automation to reduce manual ops and improve reliability.
  • Mentor teammates on SRE best practices and standards.

Skills

BSc in CS
5+ years SRE
Unix/Linux
Cloud platforms
IaC
Kubernetes
Monitoring tools
Networking

Education

Bachelor's degree

Tools

Terraform
Ansible
Docker
Kubernetes
Datadog
Prometheus

Job description

Chess.com is seeking a Site Reliability Engineer to ensure stability, performance, and scalability of our global gaming platform infrastructure. You will support millions of concurrent users, own on-call rotations, and drive incident response to maintain high availability.

You’ll collaborate across engineering teams to implement infrastructure-as-code, lead hybrid cloud migrations, and optimize monitoring, security, and capacity planning for a robust, scalable system.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

chess • United States

Remote
USD 150,000 - 190,000
Senior Site Reliability Engineer - Scale Global Platforms
Senior Site Reliability Engineer - Scale Global Platforms

Penn Interactive Ventures • California (MO)

Hybrid
USD 105,000 - 140,000
Competitive compensation
Relaxed work environment
Education reimbursements
Senior Site Reliability Engineer: Scale & Resilience
Senior Site Reliability Engineer: Scale & Resilience

Scientific Games • Alpharetta (GA)

On-site
USD 120,000 - 180,000
Remote Senior Site Reliability Engineer — AI‑Driven Infra
Remote Senior Site Reliability Engineer — AI‑Driven Infra

Precisely • Atlanta (GA)

On-site
USD 150,000 - 190,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer — Remote Production Reliability
Senior Site Reliability Engineer — Remote Production Reliability

Fingerprint • Chicago (IL)

Remote
USD 152,000 - 205,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Senior Incident Command & Reliability Engineer
Senior Incident Command & Reliability Engineer

IBM • Boston (MA)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Remote Site Reliability Engineer II — Scale & Reliability
Remote Site Reliability Engineer II — Scale & Reliability

Mrsool • United States

Remote
USD 120,000 - 180,000
Remote work options
Competitive compensation
Learning stipend