Site Reliability Engineer

iSoftStone

Kuala Lumpur

On-site

MYR 90,000 - 150,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

iSoftStone seeks an experienced SRE/DevOps engineer to design, deploy and maintain overseas game platforms. You will ensure high availability and scalable services, focusing on account and data storage, while monitoring production environments and resolving performance issues.

You'll also build observability solutions, automate tasks, and contribute to robust incident RCA practices. Strong English and Mandarin communication are essential for global collaboration, with a background in Unix/Linux,

Qualifications

  • 1+ year of hands-on experience in SRE, DevOps, Cloud Engineering, or Platform Operations.
  • Experience in the gaming industry is a strong advantage.
  • Strong knowledge of Unix/Linux operating systems; basic troubleshooting.
  • Hands-on with Kubernetes and cloud platforms such as AWS or GCP.
  • Experience with MySQL, Redis or related database technologies.
  • Good English and Mandarin communication skills.

Responsibilities

  • Participate in architecture design, deployment, and maintenance of overseas game application platforms.
  • Ensure high availability, reliability, and scalability of game platform services, especially account and data storage.
  • Monitor and maintain production environments; proactively resolve performance and reliability issues.
  • Design, optimize, and maintain monitoring and observability solutions for better visibility.
  • Support incident troubleshooting, RCA, and preventive measures to minimize disruptions.
  • Automate routine tasks and continuously improve deployment, monitoring, and maintenance processes.
  • Maintain technical documentation, operational procedures, and troubleshooting guidelines.

Skills

Kubernetes
Unix/Linux
Shell scripting
Python
Go
Public cloud AWS/GCP
Monitoring & observability
Prometheus
Grafana
ELK/EFK
MySQL/Redis
CI/CD automation
RCA & incident response
English communication
Mandarin communication

Education

Bachelor's degree in Computer Science / Information Technology / Software Engineering

Tools

Container orchestration tools
Cloud platforms

Job description

A leading global technology group, renowned for its extensive ecosystem of digital services and platforms. With a strong presence in cloud computing, mobile gaming, social media, and enterprise solutions, the organization supports millions of users and businesses worldwide. It emphasizes innovation, scalability, and security, making it a key player in driving digital transformation across various industries.

  • Participate in the architecture design, deployment, and maintenance of overseas game application platforms.
  • Ensure the high availability, reliability, and scalability of game platform services, particularly account and data storage services.
  • Monitor and maintain production environments, proactively identifying and resolving performance, availability, and reliability issues.
  • Design, optimize, and maintain monitoring and observability solutions to improve platform visibility and operational efficiency.
  • Support incident troubleshooting, root cause analysis (RCA), and preventive measures to minimize service disruptions.
  • Automate routine operational tasks and continuously improve deployment, monitoring, and maintenance processes.
  • Maintain technical documentation, operational procedures, and troubleshooting guidelines.
  • Bachelor's degree or above in Computer Science, Software Engineering, Information Technology, or a related technical field.
  • Minimum 1 year of hands-on experience in SRE, DevOps, Cloud Engineering, or Platform Operations. Experience in the gaming industry is a strong advantage.
  • Strong knowledge of Unix/Linux operating systems, with hands-on experience in system troubleshooting.
  • Practical experience with Shell and/or Python scripting for automation and operational tasks.
  • Hands-on experience managing public cloud platforms, such as AWS or GCP.
  • Solid experience with Kubernetes (K8s) and its ecosystem, including containerized application deployment and operations.
  • Working experience with MySQL, Redis, or related database technologies.
  • Understanding of monitoring and observability, with experience using tools such as Prometheus, Grafana, ELK/EFK, or similar technologies is an advantage.
  • Good English and Chinese (Mandarin) communication skills are required, as the role involves collaboration with global team.
  • Software development experience in Python, Go, or other programming languages is a plus point.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Reliability Engineer (Gaming) (3rd Party Contract -1 Year Renewable)
Platform Reliability Engineer (Gaming) (3rd Party Contract -1 Year Renewable)

Tencent • Kuala Lumpur

On-site
MYR 60,000 - 80,000
Global Game Platform SRE: Reliability & Automation
Global Game Platform SRE: Reliability & Automation

iSoftStone • Kuala Lumpur

On-site
MYR 90,000 - 150,000
Global Platform Reliability Engineer - Gaming
Global Platform Reliability Engineer - Gaming

Tencent • Kuala Lumpur

On-site
MYR 60,000 - 80,000
IEG - SRE (3rd Party Contract -1 Year Renewable)
IEG - SRE (3rd Party Contract -1 Year Renewable)

Tencent • Kuala Lumpur

On-site
MYR 150,000 - 230,000
Cloud Operations Engineer (Platform Reliability / NOC)
Cloud Operations Engineer (Platform Reliability / NOC)

Agensi Pekerjaan Genie Hunt Talent • Petaling Jaya

On-site
MYR 60,000 - 80,000
Site Reliability Engineer
Site Reliability Engineer

Confidential Jobs • Kuala Lumpur

On-site
MYR 150,000 - 230,000
Devops Engineer
Devops Engineer

Lightspeed Studios • Kuala Lumpur

On-site
MYR 60,000 - 90,000
SRE Engineer (DevOps)
SRE Engineer (DevOps)

Ant International • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Technical Service Engineer Program Manager
Technical Service Engineer Program Manager

CSI Interfusion Inc. • Malaysia

On-site
MYR 180,000 - 280,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

SGS (Malaysia) Sdn Bhd • Kuching

On-site
MYR 120,000 - 180,000