Senior Site Reliability Engineer

Accelbyte, Inc.

Daerah Istimewa Yogyakarta

On-site

IDR 360,000,000 - 600,000,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

AccelByte, Inc. is building a 24x7 SRE team in Indonesia to maintain high reliability of our game services. You will design, implement, and operate scalable Kubernetes deployments, automate operations with Terraform and CI/CD, and build Go/Python code for core platform components.

The role focuses on incident response, engineering excellence, and cross-team collaboration across time zones. You will work with CNCF tools to monitor systems, improve stability, and implement AIOps solutions, while

Qualifications

  • 5+ years Cloud Engineering or DevOps experience with AWS and Kubernetes
  • Deep knowledge of cloud providers, IaC, and GitOps
  • Experience building infrastructure as code with Terraform and package management (Helm)
  • Strong Go and Python coding for automation and tooling
  • Proven ability to design scalable, reliable backend services and deployment frameworks
  • Experience with monitoring, alerting, and incident response (Prometheus, Grafana, CloudWatch)
  • Familiarity with security best practices and cross-team collaboration

Responsibilities

  • Design, build and maintain high-performance backend services and tooling for operations automation
  • Architect and maintain scalable deployment frameworks and tooling
  • Build and operate deployments with Kubernetes and CNCF projects
  • Provide secure, cost-effective cloud platform and infrastructure
  • Create monitoring for health, outages, and root cause analyses
  • Develop automation to support product teams and operational excellence
  • Document infrastructure and runbooks; collaborate with clients and stakeholders
  • Develop AIOps solutions to improve efficiency and predictive maintenance
  • Write clean, tested Go/Python code for core platform components

Skills

Cloud engineering
Kubernetes
AWS
GitOps
Terraform
Go
Python
CI/CD
Security
Grafana
Prometheus
CloudWatch
PostgreSQL
MongoDB
ElasticSearch
Helm
Docker
Git

Education

CS Degree

Tools

Docker
Kubernetes
git
Redis
MongoDB
PostgreSQL
ElasticSearch
GitLab CI
Nexus
SonarQube
Terraform
Helm
Prometheus
ELK/EFK
Grafana
CloudWatch

Job description

POSITION SUMMARY :

AccelByte is building a 24x7 operations team for AAA multiplayer video games. In this position, we need a driven Site Reliability Engineer who can actively participate in the day-to-day combat by maintaining high reliability of our service and drive prioritization in fixing what may be broken today, as well as able to envision, design, and implement processes and technologies to improve the ability to identify, isolate, correlate, and mitigate service impacting problems in the system. The Site Reliability Engineer must also know some coding to automate routine tasks in service metrics gathering, correlating, organizing, and presenting, in addition to detail and in-depth root cause analysis

ESSENTIAL FUNCTIONS/RESPONSIBILITIES:
  • Design, build and maintain high-performance backend services, tools, and control planes to automate operations and improve system reliability.
  • Architect, implement and maintain a highly scalable deployment framework and tooling that improves our products' stability, reliability, and availability.
  • Build and run service deployment using K8s and other CNCF projects
  • Provide a secure, high-scalable, and cost-effective cloud platform
  • Construct and build effective systems to monitor the health of our system/applications, and to handle outages
  • Solve problems occurring in all our environments and create solutions to prevent them from happening again
  • Produce automation and innovative tools to assist the product development teams and to deliver operational excellence
  • Create and maintain infrastructure-related documentation and SRE runbooks
  • Collaborate with other stakeholders to provide cost-effective, operational excellence, and performance-efficient infrastructure solutions to improve our products.
  • Identify technology, process gaps, and opportunities for improvement
  • Liaise, communicate, and work directly with our clients
  • Perform any other design-related duties as required
  • Envision, design, and implement AIOps solutions to enhance operational efficiency and predictive maintenance.
  • Write clean, maintainable, and well-tested code (primarily in Go/Python) for core platform components.
QUALIFICATIONS/EXPERIENCE REQUIRED
  • 5+ years Cloud Engineering or DevOps experience with AWS, 2+ years Kubernetes, Certification in AWS preferred
  • Degree in Computer Science or equivalent experience
  • Deep knowledge of cloud service providers and best practices around implementation and configuration, preferably managing AWS and Kubernetes
  • Familiarity with infrastructure management and operations lifecycle concepts and ecosystem, deep understanding of IaC and GitOps
  • Proven track record of building infrastructure as code (Terraform is a must), configuration management, and package manager (eg: Helm Chart)
  • Experience in delivering products against a plan in a fast-paced, multi-disciplined, and often ambiguous environment
  • Experience working independently to design, plan, and execute technical projects
  • Demonstrated deep knowledge of technical program management and engineering best practices
  • Innovative thinking balanced with a strong customer and quality and cost efficiency focus
  • Comfort and experience with cross-organizational communication; excellent written and verbal communication skills
  • Working experience with some of the following technologies and tools: Docker, Kubernetes, git, Redis, MongoDB, PostgreSQL, ElasticSearch, GitLab CI, Nexus, SonarQube, Terraform, Helm, Prometheus, ELK/EFK, Grafana, CloudWatch
  • Solid security best practices
  • Strong proficiency in Go, including the ability to conduct high-quality code reviews. Experience with Python and Bash is also required.
  • Keen problem-solving skills with the ability to work under pressure (during a production event)
  • Flexibility in working with people with different timezones
  • Experience with AIOps and building/harnessing AI tools to automate and optimize operational tasks.
QUALIFICATIONS/EXPERIENCE PREFERRED
  • Previous experience working in the game industry
  • Working experience with one or more of the following: Emissary, Linkerd, Istio, Nomad, Kafka, Flux, ArgoCD, GitOps, DevSecOps
  • Familiar with web services patterns/architectures, e.g. REST, SOAP, etc.
  • Experience working with auto-scaling workloads both in containers and VMs
  • Experience with other cloud technologies and infrastructure: GCP, Azure
  • Experience with Confluence, Jira, and BitBucket
  • IT standards, methodologies, Cryptographic key management regulations, and audit experience would be asset(s).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

AccelByte • Kota Yogyakarta

Hybrid
IDR 420,000,000 - 700,000,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
IDR 167,400,000 - 279,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
Senior Site Reliability Engineer - Gaming Cloud & AIOps
Senior Site Reliability Engineer - Gaming Cloud & AIOps

AccelByte • Kota Yogyakarta

Hybrid
IDR 420,000,000 - 700,000,000
24/7 Game SRE — Cloud Infra & Reliability Lead
24/7 Game SRE — Cloud Infra & Reliability Lead

AccelByte • Sleman

On-site
Senior Site Reliability Engineer: AIOps & Cloud Automation
Senior Site Reliability Engineer: AIOps & Cloud Automation

Accelbyte, Inc. • Daerah Istimewa Yogyakarta

On-site
IDR 360,000,000 - 600,000,000
Senior Backend Engineer
Senior Backend Engineer

ACCELBYTE • Sleman

On-site
IDR 525,210,084 - 1,050,420,168
Site Reliability Engineer
Site Reliability Engineer

AlloFresh • Jakarta Pusat

On-site
IDR 450,000,000 - 750,000,000
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

PT Media Indonusa (Jakarta) • Jakarta Utara

On-site
IDR 420,000,000 - 540,000,000
Site Reliability Engineer / DevOps
Site Reliability Engineer / DevOps

Catalyst Tech • Daerah Khusus Ibukota Jakarta

On-site
IDR 720,720,000 - 1,081,082,000