Site Reliability Engineer (SRE)

AccelByte

Sleman

On-site

IDR 267,840,000 - 468,720,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AccelByte is building a 24x7 operations team for AAA multiplayer video games. We seek a driven Site Reliability Engineer who can maintain high reliability, automate tasks, and drive incident response and root cause analysis.

You will design and maintain infrastructure with Kubernetes, cloud platforms, and IaC, plus collaborate with product teams and clients to ensure cost-effective, scalable systems. Fluency in English and willingness to work shifts are required.

Qualifications

  • 2+ years Linux administration.
  • Degree in Computer Science or equivalent.
  • Experience designing, managing and running large-scale cloud apps.
  • Experience with monitoring systems and strategies.
  • Strong troubleshooting and performance skills.
  • Strong foundation in distributed systems.
  • Cloud experience (AWS/GCP) preferred.
  • Experience with containerization (Docker/Kubernetes).
  • IaC and configuration management (Terraform, Helm).
  • Automation/CICD and GitOps tools (Jenkins, GitLab, Flux, ArgoCD).
  • Experience with monitoring/alerting tools (Prometheus, Grafana).
  • Greenfield environment experience and scripting (Bash/Python/Go).
  • Ability to work with clients under tight deadlines.
  • Good communication, English fluency.
  • Willing to work on a 24/7 shift.

Responsibilities

  • Design, implement, and maintain infrastructure for applications.
  • Build and run service deployments using Kubernetes and CNCF projects.
  • Provide a secure, scalable cloud platform.
  • Monitor system health and handle outages.
  • Develop automation and tools to improve operations.
  • Create and maintain infrastructure documentation and runbooks.
  • Collaborate with stakeholders to deliver cost-efficient infrastructure.

Skills

Linux admin
Cloud (AWS/GCP)
Docker & Kubernetes
Terraform
GitOps CI/CD
Scripting (Python/Bash/Go)
Monitoring & alerting
English fluent
Client communication
24/7 shift

Education

Bachelor’s degree in CS or equivalent

Tools

Terraform
Jenkins
GitLab
GitHub
Flux
ArgoCD
Prometheus
Grafana
ELK/EFK
Datadog
PagerDuty

Job description

POSITION SUMMARY

AccelByte is building a 24x7 operations team for AAA multiplayer video games. In this position, we need a driven Site Reliability Engineer who can actively participate in the day‑to‑day combat by maintaining high reliability of our service and driving prioritization in fixing what may be broken today, as well as envision, design and implement processes and technologies to improve the ability to identify, isolate, correlate and mitigate service‑impacting problems in the system. The Site Reliability Engineer must also know some coding to automate routine tasks in service metrics gathering, correlating, organizing, and presenting, in addition to detail and in‑depth root cause analysis.

ESSENTIAL FUNCTIONS / RESPONSIBILITIES
  • Design, implement, and maintain infrastructure for applications
  • Build and run service deployment using K8s and other CNCF projects
  • Provide a secure, high‑scalable, and cost‑effective cloud platform
  • Construct and build effective systems to monitor the health of our system/applications, and to handle outages
  • Solve problems occurring in all our environments and create solutions to prevent them from happening again
  • Produce automation and innovative tools to assist the product development teams and to deliver operational excellence
  • Create and maintain infrastructure related documentations and SRE runbooks
  • Collaborate with other stakeholders to provide cost‑effective, operational excellence, and performance efficient infrastructure solutions to improve our products.
  • Identify technology, process gaps, and opportunities for improvement
  • Liaise, communicate, and work directly with our client.
QUALIFICATIONS / EXPERIENCE REQUIRED
  • 2+ years Linux administration
  • Degree in Computer Science or equivalent experience
  • Prior experience helping design, manage and run large scale applications in the cloud
  • Experience with monitoring systems and strategies (System Admin)
  • Solid performance and troubleshooting skills
  • Solid foundation on distributed system
  • Robust knowledge and experience in cloud computing of at least one cloud provider (preferred AWS/GCP)
  • Experience with containerization principles and frameworks such as Docker, Container, Kubernetes, etc
  • Proven track record of building infrastructure as code (Terraform is must), configuration management, and package manager (eg: Helm Chart)
  • Proven experience with automation, CICD, and GitOps tools such as Jenkins, GitLab, GitHub, Flux, and/or ArgoCD
  • Experience with monitoring and alerting tools such as Prometheus, Grafana, ELK/EFK, Splunk, Datadog, OpsGenie, PagerDuty, etc
  • Experience within a greenfield environment, building infrastructure from scratch
  • Software development and scripting experience with Bash, Python, and/or Golang
  • Ability to work with clients on tight deadlines and fluid requirements
  • Good communication skill (escalation, explaining the incident)
  • Fluent in English both spoken and written
  • Willing to work on shift (24/7)
QUALIFICATIONS / EXPERIENCE PREFERRED
  • Contribute to open source projects and participate in technical communities
  • Experience working for or with AAA game studios
  • JVM tuning and troubleshooting
  • Experience with web services
  • Experience in Networking, Security, or Storage
  • Experience managing SQL and NoSQL databases
  • Familiar with Perforce version control
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

AccelByte • Kota Yogyakarta

On-site
IDR 2,144,389,000 - 3,216,583,000
24/7 Game SRE — Cloud Infra & Reliability Lead
24/7 Game SRE — Cloud Infra & Reliability Lead

AccelByte • Sleman

On-site
Senior SRE: Cloud-Native Reliability for AAA Games
Senior SRE: Cloud-Native Reliability for AAA Games

AccelByte • Kota Yogyakarta

On-site
IDR 2,144,389,000 - 3,216,583,000
Site Reliability Engineer
Site Reliability Engineer

StraitsX • Daerah Khusus Ibukota Jakarta

On-site
Site Reliability Engineer
Site Reliability Engineer

PARTECH PARTNERS • Daerah Khusus Ibukota Jakarta

On-site
IDR 272,380,000 - 453,968,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

StraitsX • Indonesia

On-site
IDR 300,000,000 - 540,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Quiet Capital • Jakarta Pusat

On-site
IDR 350,000,000 - 700,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

StraitsX • Jakarta Pusat

On-site
IDR 600,000,000 - 840,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BookCabin • Jakarta Pusat

On-site
IDR 250,000,000 - 420,000,000