Senior Site Reliability Engineer (Night Shift 5PM-2AM)

Razer Inc.

Kuala Lumpur

On-site

MYR 180,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Razer Inc. is seeking a Senior Site Reliability Engineer to join our global infrastructure team. You will design, implement, and operate scalable cloud systems on AWS, focusing on reliability, observability, and automation.

You will collaborate with developers and security to deliver resilient services, mentor junior engineers, and drive incident response and uptime improvements. This role requires hands-on AWS, IaC, and scripting across Linux and container environments.

Qualifications

  • Bachelor’s degree in Computer Science, Software Engineering, IT, or related.
  • Minimum 3 years in SRE/DevOps, cloud infra, or sysadmin roles.
  • Hands-on AWS: EC2, Lambda, ECS, EKS, Auto Scaling, VPC, Route 53, S3, RDS, ElastiCache, SQS, SES.

Responsibilities

  • Design, implement, and maintain IaC using Terraform and CloudFormation across multi-account AWS.
  • Collaborate with developers, architects, and DevOps to build scalable, secure cloud infra.
  • Lead architecture discussions for reliability, security, performance, and observability.
  • Implement monitoring/alerting (CloudWatch, Prometheus, Datadog) and drive incident responses.

Skills

Troubleshooting
Incident management
Mentoring
Monitoring/Observability
SRE best practices

Education

Bachelor’s degree in Computer Science or related

Tools

AWS Cloud Services
Terraform
CloudFormation
Python/Node.js/Bash
Linux/Windows/containers
Datadog/CloudWatch/ELK

Job description

Joining Razer will place you on a global mission to revolutionize the way the world games. Razer is a place to do great work , offering you the opportunity to make an impact globally while working across a global team located across 5 continents. Razer is also a great place to work, providing you the unique, gamer-centric #LifeAtRazer experience that will put you in an accelerated growth, both personally and professionally.

Job Responsibilities :

We are seeking a skilled and driven Senior Site Reliability Engineer (SRE) to join our growing infrastructure and platform engineering team. The ideal candidate will have hands-on experience in Amazon Web Services (AWS), strong troubleshooting capabilities, and a passion for building scalable, observable, and resilient systems using modern Infrastructure as Code (IaC) and automation tools.

REQUIREMENTS:

  • Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related field.
  • Minimum 3 years of experience in SRE, DevOps, cloud infrastructure, or system administration roles.
  • Hands-on expertise with AWS Cloud Services, including:
  • Compute & Containerization: EC2, Lambda, ECS, EKS, Auto Scaling
  • Networking: Load Balancers, VPC, Route 53, Security Groups, Firewalls
  • Storage & Databases: RDS, ElastiCache, Athena, S3
  • Messaging: SQS, SES
  • Deep understanding of Infrastructure as Code (IaC) tools such as Terraform and CloudFormation.
  • Proficiency in at least one programming/scripting language: Python, Node.js, Bash, Ruby, or related.
  • Experience operating and troubleshooting across Linux, Windows, and container-based environments.
  • Strong understanding of distributed systems, cloud networking (routers, switches), firewalls, DNS, and HTTP/TLS.
  • Experience implementing monitoring and alerting systems and working with incident management processes.
  • Experience with Zero Downtime Deployments, blue/green or canary deployments.
  • Familiarity with cost optimization and right-sizing AWS resources.
  • Exposure to multi-region, multi-account AWS architecture.
  • Understanding of API gateway, or edge networking (e.g., Akamai, CloudFront).
Support from 5:00PM to 2:00AM (UTC+8) shift to ensure continuous of SRE coverage.

JOB DESCRIPTION:

  • Design, implement, and maintain Infrastructure as Code (IaC) solutions using Terraform and/or CloudFormation across multi-account AWS environments.
  • Collaborate with developers, architects, and DevOps teams to build scalable, secure, and observable cloud infrastructure.
  • Lead and participate in architecture design sessions, focusing on system reliability, scalability, security, and performance.
  • Implement and manage robust monitoring, alerting, and observability solutions (e.g., CloudWatch, Prometheus, ELK, Datadog).
  • Set and monitor Key Performance Indicators (KPIs) for system uptime, latency, throughput, and overall reliability.
  • Drive incident response processes, including coordination, triaging, resolution, documentation, and post-incident reviews (PIRs).
  • Supervise and mentor junior SREs and infrastructure engineers, fostering knowledge-sharing and team growth.
  • Collaborate across development, operations, and security teams to ensure secure and compliant deployments.
  • Automate manual tasks and workflows through scripting and tooling (Python, Node.js, Bash, Ruby, JSON/YAML).
  • Troubleshoot complex infrastructure issues across Linux, Windows, Docker, and cloud-native environments.
  • Provide IaC and CI/CD best practices to ensure repeatability, scalability, and compliance across all environments.
  • Provide on-call support, participate in incident rotations, and lead technical investigations during outages or degradations.
  • Strong understanding and experience for Disaster Recovery (DR).
  • Provide support and solution handling to incident and tickets assigned.
Pre-Requisites :

Razer is proud to be an Equal Opportunity Employer. We believe that diverse teams drive better ideas, better products, and a stronger culture. We are committed to providing an inclusive, respectful, and fair workplace for every employee across all the countries we operate in. We do not discriminate on the basis of race, ethnicity, colour, nationality, ancestry, religion, age, sex, sexual orientation, gender identity or expression, disability, marital status, or any other characteristic protected under local laws. Where needed, we provide reasonable accommodations - including for disability or religious practices - to ensure every team member can perform and contribute at their best.

Are you game?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (Night Shift 5PM-2AM)
Senior Site Reliability Engineer (Night Shift 5PM-2AM)

Razer • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Senior Site Reliability Engineer (Night Shift 5PM-2AM)
Senior Site Reliability Engineer (Night Shift 5PM-2AM)

Razer Inc. • Malaysia

On-site
MYR 180,000 - 320,000
Senior Site Reliability Engineer (Night Shift 5PM-2AM)
Senior Site Reliability Engineer (Night Shift 5PM-2AM)

MOL Accessportal Sdn Bhd • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Senior Cloud Systems Administrator
Senior Cloud Systems Administrator

Razer Inc. • Shah Alam

On-site
MYR 100,000 - 160,000
Senior Cloud Systems Administrator
Senior Cloud Systems Administrator

Razer • Shah Alam

On-site
MYR 120,000 - 180,000
Product Engineer
Product Engineer

Razer • Shah Alam

On-site
MYR 120,000 - 180,000
Senior Automation QA Engineer
Senior Automation QA Engineer

JobCubby • Shah Alam

On-site
MYR 120,000 - 180,000
Senior Cloud Systems Administrator
Senior Cloud Systems Administrator

Razer • Shah Alam

On-site
MYR 120,000 - 180,000
Global Senior SRE: AWS, IaC & Resilience Lead
Global Senior SRE: AWS, IaC & Resilience Lead

Razer Inc. • Kuala Lumpur

On-site
MYR 180,000 - 260,000
Security Specialist
Security Specialist

Razer Merchant Services Sdn. Bhd. • Shah Alam

On-site
MYR 120,000 - 180,000