Senior Site Reliability Engineer

Engg

Kuala Lumpur

On-site

MYR 180,000 - 300,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Gen in Kuala Lumpur seeks a Senior Site Reliability Engineer to design, operate, and evolve the infrastructure powering our MoneyLion-backed, customer-focused platform.

You will work across cloud infrastructure, Kubernetes, networking, CI/CD, observability, reliability, and automation. You will mentor engineers and drive engineering best practices within the SRE organization.

Qualifications

  • Strong software and systems engineering fundamentals in an enterprise environment.
  • Hands-on AWS and cloud-native infrastructure experience.
  • Kubernetes, preferably Amazon EKS, with architecture, networking and operations knowledge.
  • Terraform IaC and automated workflows.
  • CI/CD and GitOps using GitHub Actions or similar.
  • Observability/monitoring with Datadog, Prometheus and Grafana.
  • Networking concepts including DNS, load balancing, proxies and firewalls.
  • Python, Go or Bash for automation.
  • Experience troubleshooting distributed systems and incidents across layers.
  • SRE practices: SLOs, SLIs, incident mgmt, capacity planning and disaster recovery.
  • Experience evaluating new tech, automation, and AI-assisted engineering tools.

Responsibilities

  • Design, build, operate, and improve highly available infrastructure for business-critical apps.
  • Build and operate cloud infrastructure and Kubernetes platforms with emphasis on reliability and security.
  • Develop and improve IaC, CI/CD platforms, and GitOps workflows for safe delivery.
  • Partner with engineering teams to design reliable, scalable architectures.
  • Improve observability and monitoring across applications and infrastructure.
  • Lead on-call rotations, RCAs, and drive preventive improvements.

Skills

AWS
Kubernetes
Terraform
GitOps
CI/CD
Python
Go
Bash
Datadog
Prometheus
Grafana
Networking
SRE practices

Job description

About Gen:

Gen is a global company dedicated to powering Digital Freedom through its trusted consumer brands including Norton, Avast, LifeLock, MoneyLion and more. Our combined heritage is rooted in financial empowerment and cyber safety for the first digital generations, and today we deliver award-winning cybersecurity, online privacy, identity protection and financial wellness solutions to nearly 500 million users in more than 150 countries. Together, we share a collective passion and vision to protect consumers and help them grow, manage and secure their digital and financial lives. We're always looking for smart, fearless and high-impact talent who see AI as a teammate - leveraging it to move faster and deliver meaningful results. When you're part of Gen, you'll have the flexibility, tools and support to do your best work and grow your career - from flexible working options and time off to competitive pay, benefits and well-being programs. At Gen, we are scrappy and relentlessly customer driven. We create room for healthy debate, experimentation and continuous learning, and we seek out people with different experiences, identities and ideas to join our team. You'll work with people who back each other, respect each other and understand that our differences are a competitive advantage. If this sounds like you, we'd love you to be part of Gen.

About the Role

The Kuala Lumpur office is the technology powerhouse of MoneyLion. We pride ourselves on innovative initiatives and thrive in a fast paced and challenging environment. Join our multicultural team of visionaries and industry rebels in disrupting the traditional finance industry! As a Senior Site Reliability Engineer, you will have the opportunity to build, operate, and evolve the infrastructure and platforms that power MoneyLion's business-critical applications. You will work across cloud infrastructure, Kubernetes, networking, CI/CD, observability, reliability, and automation to ensure our platforms remain highly available, scalable, secure, and efficient as the business continues to grow. You will partner closely with Software Engineers, Platform Engineers, Security teams, and other technology teams across MoneyLion to improve reliability and developer experience. As a senior member of the SRE team, you will also lead infrastructure initiatives, drive automation and platform improvements, mentor engineers, and help establish engineering and operational best practices across the organization.

Key Responsibilities
  • Design, build, operate, and continuously improve highly available and scalable infrastructure supporting MoneyLion's business-critical applications.
  • Build and operate cloud infrastructure and Kubernetes platforms, with a strong focus on reliability, scalability, security, and operational efficiency.
  • Develop and improve Infrastructure as Code, CI/CD platforms, GitOps workflows, and developer self-service capabilities to enable engineering teams to deliver software safely and efficiently.
  • Partner with engineering teams to design reliable and scalable architectures for new and existing applications and services.
  • Improve observability across applications and infrastructure through effective monitoring, alerting, logging, tracing, and automation.
  • Participate in the SRE on-call rotation, lead response to critical production incidents, perform root cause analysis, and drive corrective and preventive improvements.
  • Identify opportunities to reduce operational toil through automation, AI-assisted operations, and improvements to engineering workflows.
  • Lead infrastructure modernization, platform migration, and technology standardization initiatives across MoneyLion.
  • Improve the resilience of MoneyLion's platforms through capacity planning, performance engineering, disaster recovery, business continuity, and reliability testing.
  • Manage and optimize cloud infrastructure usage and costs while maintaining appropriate levels of performance, reliability, and scalability.
  • Collaborate closely with Security and Engineering teams to ensure infrastructure and platforms follow security, compliance, and operational best practices.
  • Mentor junior engineers, conduct technical and code reviews, and help establish engineering standards and best practices across the SRE and wider engineering organization.
About You
  • Strong software and systems engineering fundamentals, with experience building and operating large-scale production systems in an enterprise environment.
  • Strong hands-on experience with AWS and cloud-native infrastructure, including designing and operating highly available production environments.
  • Demonstrated practical experience with Kubernetes, preferably Amazon EKS, with a strong understanding of Kubernetes architecture, networking, workloads, security, and operations.
  • Strong experience with Infrastructure as Code, preferably Terraform, and experience managing infrastructure through version-controlled and automated workflows.
  • Experience with CI/CD and GitOps practices using platforms such as GitHub Actions or similar technologies.
  • Hands-on experience with observability and monitoring platforms such as Datadog, Prometheus, Grafana, or similar tools.
  • Good understanding of networking concepts including DNS, load balancing, proxies, firewalls, ingress, and cloud networking.
  • Proficiency in scripting or programming languages such as Python, Go, or Bash, with the ability to build automation and operational tooling.
  • Experience troubleshooting complex distributed systems and resolving production incidents across applications, infrastructure, and networking layers.
  • Strong understanding of Site Reliability Engineering practices, including SLOs, SLIs, incident management, capacity planning, disaster recovery, and reducing operational toil.
  • Experience evaluating and adopting new technologies, automation, and AI-assisted engineering tools to improve platform reliability and engineering productivity.
  • Stron
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Apply • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Gen • Kuala Lumpur

On-site
MYR 220,000 - 320,000
Senior SRE: Cloud, Kubernetes & Automation Leader
Senior SRE: Cloud, Kubernetes & Automation Leader

Apply • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Senior SRE: Cloud Native, Kubernetes, Automation Lead
Senior SRE: Cloud Native, Kubernetes, Automation Lead

Gen • Kuala Lumpur

On-site
MYR 220,000 - 320,000
Lead Engineer (Backend) - MoneyLion
Lead Engineer (Backend) - MoneyLion

Gen • Kuala Lumpur

On-site
MYR 80,000 - 120,000
Data Engineer, Intern - MoneyLion
Data Engineer, Intern - MoneyLion

Gen • Kuala Lumpur

On-site
Backend Engineer II - MoneyLion
Backend Engineer II - MoneyLion

Gen • Kuala Lumpur

On-site
MYR 50,000 - 70,000
Flexible working options
Competitive pay
Benefits and well-being programs
Engineering Manager (MLOps) - MoneyLion
Engineering Manager (MLOps) - MoneyLion

Gen • Kuala Lumpur

On-site
MYR 100,000 - 150,000
Flexible working options
Competitive pay
Well-being programs
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Gen • Kuala Lumpur

On-site
MYR 80,000 - 120,000
Flexible working options
Competitive pay and benefits
Training and development programs
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

SGS (Malaysia) Sdn Bhd • Kuching

On-site
MYR 120,000 - 180,000