Site Reliability Engineer (4024)

Uncover

Kuala Lumpur

On-site

MYR 180,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GBG is seeking a Site Reliability Engineer to build and operate the reliability and observability stack for its Global Fraud Solutions platform. You will work across deployment pipelines, cloud infra, monitoring, and incident management to meet high-availability SLAs for real-time fraud detection.

You will design on-call processes, implement IaC using Terraform, manage CI/CD pipelines, and ensure PCI DSS/ISO/GDPR compliance in collaboration with InfoSec.

Qualifications

  • 5+ years of hands-on experience in Site Reliability/DevOps, supporting real-time processing systems.
  • Experience with cloud platforms (AWS preferred; Azure/GCP acceptable) and orchestration.
  • Strong expertise in observability: metrics, logging, tracing, and alerting.
  • Knowledge of PCI DSS, ISO 27001, SOC 1/2/3, GDPR/PDPA and security best practices.

Responsibilities

  • Design and operate the SRE practice for Managed offerings, including on-call processes and incident reviews.
  • Build and maintain observability infrastructure: centralized logging, dashboards, tracing, and alerting.
  • Define and track SLOs and error budgets for real-time pipelines with high TPS.
  • Manage IaC provisioning for AWS/Azure and on-prem environments using Terraform/Helm.
  • Implement CI/CD pipelines and ensure security/compliance readiness across hosted services.
  • Drive platform resilience: HA, auto-scaling, DR, and chaos engineering practices.
  • Collaborate with engineering teams to embed reliability and DevSecOps across the SDLC.

Skills

Site Reliability
DevOps
Cloud platforms
Docker
Kubernetes
Terraform
CI/CD
Security & Compliance
Python
Incident response

Tools

Terraform
Docker
Kubernetes
Jenkins
GitHub Actions
Prometheus
Grafana
ELK/OpenSearch
Jaeger
HashiCorp Vault
AWS

Job description

Enabling safe and rewarding digital lives for genuine people, everywhere. We make it our mission to ensure more genuine people have digital access to opportunities, and businesses have access to more genuine people. Our technology draws on diverse and reliable data to create a single point of truth for identity and address verification. With over 30 years of experience behind us our team and technology are focused on enabling safe and rewarding digital lives for everyone. Regardless of age, location or background, genuine people everywhere should be able to digitally prove who they are and where they live.

About the team and role
Global Fraud Solutions
Site Reliability Engineer

The SRE will build and operate the reliability, observability, and operational excellence infrastructure underpinning the GFS managed fraud detection platforms. You will work across deployment pipelines, cloud infrastructure, monitoring, and incident management — ensuring GBG can deliver on high availability SLAs for banking and fintech customers who depend on real-time fraud detection at scale.

What you will do
  • Design and operate the SRE practice for Managed oferings, including on-call processes, SLA frameworks, incident response playbooks, and post-incident review (PIR) processes.
  • Build and maintain observability infrastructure: centralised logging (correlation IDs), metrics dashboards, distributed tracing, and alerting for the Predator/Instinct platform stack.
  • Define and track SLOs (Service Level Objectives) and error budgets for real-time transaction processing pipelines, targeting high TPS and low round-trip latency.
  • Manage cloud infrastructure provisioning and configuration using IaC tooling (Terraform, Helm), supporting both AWS/Azure cloud deployments and on-premises customer environments.
  • Implement and maintain CI/CD pipelines for GFS solutions (Jenkins, etc.)
  • Work with Engineering teams to ensure security and compliance readiness for Managed services — including PCI DSS, ISO 27001, SOC 1/2/3, PDPA/GDPR — in close coordination with InfoSec teams.
  • Drive platform resilience improvements: high availability, auto-scaling, disaster recovery, backup/restore procedures, and chaos engineering practices.
  • Manage secrets, certificate rotation, identity/access controls (OAuth/RBAC), and vulnerability management for the hosted environment.
  • Support performance testing methodology and baseline establishment for our products.
  • Contribute to the Architecture Review Committee (ARC) with SRE and operational perspectives on technology choices.
  • Collaborate with engineering squads to embed reliability and DevSecOps practices across the SDLC.
Skills we’re looking for
  • Minimum 5 years of solid hands‑on experience in a Site Reliability, Platform Engineering, or DevOps role, ideally supporting mission-critical real-time processing systems in banking, payments, or fintech.
  • Strong proficiency with cloud platforms (AWS preferred; Azure/GCP acceptable) including networking, compute, storage, and managed services.
  • Deep expertise with containerisation and orchestration: Docker, Kubernetes (EKS/AKS/GKE), Helm, and associated tooling.
  • Infrastructure as Code experience: Terraform (required), and familiarity with Ansible or Pulumi.
  • Observability stack proficiency: Prometheus, Grafana, ELK/OpenSearch, Jaeger/Zipkin, or equivalent enterprise-grade tooling.
  • CI/CD pipeline design and management: GitHub Actions, Jenkins, ArgoCD, or equivalent.
  • Experience with security and compliance frameworks applicable to hosted financial services: PCI DSS, ISO 27001, SOC 1/2/3, GDPR/PDPA.
  • Familiarity with database reliability practices for SQL Server, PostgreSQL, and Oracle — including replication, read replicas, and failover.
  • Working knowledge of secrets management (HashiCorp Vault, AWS Secrets Manager) and zero-trust identity principles.
  • Experience supporting real-time streaming or event-driven architectures (Kafka, RisingWave, or similar) in production environments.
  • Scripting and automation proficiency: Python, Bash, or Go for operational tooling.
  • Strong sense of operational ownership and accountability — comfortable being on-call and driving incidents to resolution.
  • Excellent communication skills — able to produce clear incident reports, runbooks, and architecture documentation for both technical and executive audiences.
  • Proactive mindset: identifies reliability risks before they become incidents and champions a culture of blameless post-mortems.
  • Collaborative and effective working with software engineers, product managers, and InfoSec teams.
  • Continuous improvement orientation — always looking to reduce toil, automate repetitive tasks, and improve platform resilience.
  • Flexible and adaptable — able to support a globally distributed product with customers across multiple time zones.
To find out more

As an equal opportunity employer, we are dedicated to creating a diverse and inclusive workplace where everyone feels valued and empowered. Please inform your GBG Talent Attraction Partner if you require any reasonable adjustments to the interview process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (4024)
Site Reliability Engineer (4024)

GBG Plc • Kuala Lumpur

On-site
MYR 100,000 - 130,000
Senior Site Reliability Engineer - Real-Time FinTech
Senior Site Reliability Engineer - Real-Time FinTech

Uncover • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Ryt Bank • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Senior Site Reliability Engineer - Real-Time Fraud Platform
Senior Site Reliability Engineer - Real-Time Fraud Platform

GBG • Kuala Lumpur

On-site
MYR 60,000 - 100,000
Diverse and inclusive workplace
Employee benefits
Career development opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AirAsia rewards • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Senior Site Reliability Engineer - Real-Time Fraud Platform
Senior Site Reliability Engineer - Real-Time Fraud Platform

GBG Plc • Kuala Lumpur

On-site
MYR 100,000 - 130,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Swift Software • Kuala Lumpur

On-site
MYR 120,000 - 240,000
System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Hong Kong and Macau • Kuala Lumpur

On-site
MYR 70,000 - 90,000
System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Malaysia • Kuala Lumpur

On-site
MYR 70,000 - 110,000
High-impact team environment
Opportunities for innovation
Influence engineering culture
Azure Infrastructure & Data Engineer
Azure Infrastructure & Data Engineer

S&P Global, Inc. • Malaysia

On-site
MYR 120,000 - 240,000
Health care coverage
Flexible time off
Continuous learning opportunities
+3