Site Reliability Engineer

Aisling Group

Kuala Lumpur

On-site

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A tech eCommerce scale-up in Kuala Lumpur is seeking a Site Reliability Engineer (SRE) to ensure that services and production systems run smoothly. Responsibilities include improving the lifecycle of services, collaborating with engineering teams, and maintaining service reliability. Ideal candidates should have 5-8 years of experience in infrastructure management, proficiency in Python, Go, or Ruby, and familiarity with cloud environments. This role offers the chance to work on complex challenges and innovate in a dynamic environment.

Qualifications

  • 5-8 years of experience in provisioning environments and deploying applications.
  • Extensive experience building scalable platforms leveraging containers.
  • Experience in cloud environments such as AWS, GCP, or Azure.
  • Ability to identify and prevent outages systematically.

Responsibilities

  • Engage in and improve the lifecycle of services from design to operation.
  • Collaborate with engineering teams on infrastructure needs.
  • Maintain live services by monitoring availability and system health.
  • Scale systems sustainably through automation.

Skills

Proficient in Python, Go, or Ruby
Experience with deployment automation tools
Knowledge of continuous integration and delivery
Ability to debug and optimize code
Systematic problem-solving approach
Effective communication skills

Tools

Chef
Ansible
Terraform
Elasticsearch
SQL Azure

Job description

Kuala Lumpur, Federal Territory of Kuala Lumpur, Malaysia

Our client is a Tech Ecommerce Scale-Up that provides a single platform for customers to shop for the best price online. Not only that, they also provide data and insights to customers on latest trends and e-commerce sector.

They are looking for a Site Reliability Engineers (SREs) who are responsible for keeping all services and production systems running smoothly. SREs ensures that services have reliability, uptime appropriate to users' needs and a fast rate of improvement.

You'll have the opportunity to work on complex challenges of scale, using your experience in coding, algorithms, and analysis

RESPONSIBILITY:

  • Engage in and improve the whole lifecycle of services - from inception and design, through to deployment, operation and refinement.
  • Collaborate with engineering teams on their infrastructure needs, and advise them throughout the development lifecycle.
  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health, within our Service Level Objectives.
  • Scale systems sustainably through mechanisms like automation; evolve systems by pushing for changes that improve reliability and velocity.
  • Practice sustainable incident response and blameless post-mortems.
  • Debug production issues across services, databases and levels of the stack.
  • Design, develop and manage monitoring tools to provide performance dashboards, alerts, and collect data required to proactively identify issues and/or recommend improvements.
REQUIREMENTS:
  • 5-8 years of experience in provisioning environments, deploying applications, and maintaining infrastructures.
  • Professional experience using Python, Go, or Ruby.
  • Experience with deployment automation/configuration management tools like Chef, Ansible, Puppet, or Terraform.
  • Experience in cloud-based environment such as AWS, GCP or Azure.
  • Have extensive experience building scalable platforms leveraging containers in a production environment.
  • Added bonus if you have experience in operated distributed data storage systems at scale, especially Elasticsearch and SQL Azure.
  • Solid knowledge of continuous integration, continuous delivery, automated testing and all phases of the software development lifecycle.
  • Experience of working in an agile and multi-cultural environment across many SCRUM teams at the same time.
  • A Kaizen mindset and spirit of continuous improvement on a personal level and always up to date with the latest technology trends professionally.
  • Ability to identify problems before they happen and implement solutions that detect and prevent outages.
  • Expertise in designing, analysing and troubleshooting large-scale distributed systems.
  • Ability to debug, optimize code and automate routine tasks.
  • Systematic problem-solving approach, coupled with effective communication skills and a sense of drive.
  • Understanding of CI/CD principles, Linux fundamentals, networking concepts and IP protocols.
HOW TO APPLY:
  • If you're interested, do click apply on the button provided and attach your CV as well. For further information, feel free to speak to Ariff at +6012-9264666 or email him at ariff.w@aislingsearch.com
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Scale-Up Platform
Senior Site Reliability Engineer - Scale-Up Platform

Aisling Group • Kuala Lumpur

On-site
Site Reliability Engineer
Site Reliability Engineer

LAVU TECH SOLUTIONS SDN. BHD. • Petaling Jaya

On-site
MYR 180,000 - 300,000
Database Reliability Engineer—Scale & Automation Pro
Database Reliability Engineer—Scale & Automation Pro

Aisling Group • Kuala Lumpur

On-site
MYR 80,000 - 110,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Esker • Kuala Lumpur

Hybrid
MYR 120,000 - 180,000
Database Reliability Engineer
Database Reliability Engineer

Aisling Group • Kuala Lumpur

On-site
MYR 80,000 - 110,000
Site Reliability Engineer (SRE) Esker Asia · Kuala Lumpur, Malaysia ·
Site Reliability Engineer (SRE) Esker Asia · Kuala Lumpur, Malaysia ·

Esker, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Site Reliability Engineer: Build Reliable, Scalable Systems
Site Reliability Engineer: Build Reliable, Scalable Systems

Setel Ventures • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Leisure area with video games
Casual dress (jeans)
Pantry with coffee, tea and snacks
+2
Senior Site Reliability Engineer: Scale, Resilience & Automation
Senior Site Reliability Engineer: Scale, Resilience & Automation

LAVU TECH SOLUTIONS SDN. BHD. • Petaling Jaya

On-site
MYR 180,000 - 300,000
Immediate Opening for Site Reliability Engineer
Immediate Opening for Site Reliability Engineer

Pan Asia Group • Kuala Lumpur

On-site
MYR 140,000 - 210,000
Senior SRE: Automation, Reliability & Scale
Senior SRE: Automation, Reliability & Scale

Swift Software • Kuala Lumpur

On-site
MYR 120,000 - 240,000