Remote Site Reliability Engineer - Scalable Gaming Platform

your Jared

Northern (KY)

Hybrid

USD 150,000 - 210,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work

Job summary

Chess.com is seeking a Site Reliability Engineer to ensure the stability, performance and scalability of our global gaming platform. You will bridge development and operations, maintaining high availability for millions of users while enabling rapid feature delivery.

You will lead hybrid cloud migrations, implement infrastructure-as-code, design robust monitoring, and guide cross-functional teams in reliability best practices. This is a fully remote, full-time role with global impact.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience.
  • 5+ years of site reliability engineering, DevOps, or infrastructure engineering roles.
  • Experience managing bare-metal server infrastructure and datacenter operations.
  • Strong proficiency with UNIX/Linux operating systems and command-line administration.
  • Experience with cloud platforms (GCP, AWS, or Azure) and infrastructure-as-code tools (Terraform, CloudFormation, or similar).
  • Hands-on experience with configuration management systems (Ansible, Chef, Puppet, or similar).
  • Solid understanding of networking fundamentals, protocols (TCP/IP, HTTP/HTTPS, DNS), and network troubleshooting.
  • Experience with containerization and orchestration technologies (Docker, Kubernetes, or similar).
  • Proficiency with monitoring and observability tools (Datadog, Prometheus, Grafana, ELK stack, or similar).
  • Experience with relational and NoSQL databases, including performance optimization and scaling strategies.
  • Strong collaboration and communication skills for working effectively in a distributed team environment.
  • Demonstrated sense of ownership and accountability for system reliability and performance.

Responsibilities

  • Design and implement multi-regional resilient infrastructure capable of handling millions of concurrent sessions and transactions daily across global data centers.
  • Lead the hybrid cloud migration strategy, integrating bare-metal datacenter resources with cloud services for optimal performance and cost efficiency.
  • Own the on-call rotation and incident response procedures, ensuring rapid resolution of critical system issues and maintaining high availability SLAs.
  • Architect monitoring and alerting systems using industry-standard tools to proactively identify and resolve performance bottlenecks before they impact users.
  • Collaborate with development teams to implement infrastructure-as-code practices and establish deployment pipelines that support continuous integration and delivery.
  • Optimize system performance through capacity planning, load testing, and resource allocation across distributed computing environments.
  • Establish and maintain security protocols and risk assessment procedures for infrastructure components and data protection.
  • Partner with engineering teams to design scalable solutions for high-traffic applications and real-time processing requirements.
  • Drive automation initiatives to reduce manual operational overhead and improve system reliability through scripting and configuration management.
  • Mentor team members on SRE best practices and contribute to the development of infrastructure standards and documentation.

Skills

SRE experience
DevOps
Linux
Cloud platforms
Infrastructure as code
Configuration management
Networking
Docker/Kubernetes
Monitoring/Observability
Databases
CI/CD pipelines
Remote collaboration

Education

Bachelor's degree in CS/Engineering or related field

Tools

Terraform
CloudFormation
Ansible
Chef/Puppet
Docker
Kubernetes
Datadog
Prometheus
Grafana
ELK stack

Job description

Chess.com is seeking a Site Reliability Engineer to ensure the stability, performance and scalability of our global gaming platform. You will bridge development and operations, maintaining high availability for millions of users while enabling rapid feature delivery.

You will lead hybrid cloud migrations, implement infrastructure-as-code, design robust monitoring, and guide cross-functional teams in reliability best practices. This is a fully remote, full-time role with global impact.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

your Jared • Northern (KY)

Hybrid
USD 150,000 - 210,000
Remote work
Senior Platform Reliability Engineer - Remote
Senior Platform Reliability Engineer - Remote

Scopely • United States

Hybrid
GBP 90,000 - 140,000
Visa sponsorship
Relocation assistance
Remote-Optional Senior Site Reliability Engineer
Remote-Optional Senior Site Reliability Engineer

Multi Media, LLC • United States

On-site
USD 169,000 - 215,000
Fully Remote Optional
Health, Vision, Dental, Life Insurance
Unlimited PTO
+4
Remote Senior Data Science Lead — Chess Analytics
Remote Senior Data Science Lead — Chess Analytics

your Jared • Northern (KY)

Hybrid
USD 140,000 - 210,000
Remote Site Reliability Engineer: Scale & Resilience
Remote Site Reliability Engineer: Scale & Resilience

Bright Vision Technologies • Nashua (NH)

On-site
USD 100,000 - 180,000
Cloud-Scale Site Reliability & Infra Engineer
Cloud-Scale Site Reliability & Infra Engineer

Medal • New York (NY)

On-site
USD 120,000 - 180,000
Competitive salary and equity
Comprehensive medical, dental, vision coverage
401(k) plan
+5
Senior Site Reliability Engineer: Scale & Resilience
Senior Site Reliability Engineer: Scale & Resilience

Scientific Games • Alpharetta (GA)

On-site
USD 120,000 - 180,000
Remote Senior Site Reliability Engineer - Cloud & Automation
Remote Senior Site Reliability Engineer - Cloud & Automation

Multi Media LLC • United States

On-site
USD 169,000 - 215,000
Fully Remote
Health Insurance
Vision Insurance
+10
Senior Cloud Platform Lead Engineer - Remote
Senior Cloud Platform Lead Engineer - Remote

GN Group • United States

On-site
USD 140,000 - 200,000
10% time (learning) policy
Education stipend
Early access to SteelSeries gear
+1
Remote Senior DevOps & SRE: Multi-Cloud Reliability Lead
Remote Senior DevOps & SRE: Multi-Cloud Reliability Lead

AgileEngine, LLC. • San Francisco (CA)

On-site
USD 170,000 - 240,000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3