Associate/AVP, Observability & SRE Engineering, Technology Group

GIC Private Limited

Singapore

On-site

SGD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GIC Private Limited in Singapore seeks an Associate/AVP, Observability & SRE Engineering to lead enterprise observability and service reliability strategy across all infrastructure and applications. You will drive proactive monitoring, automation, and resilience using Datadog, Dynatrace, AWS, and Azure, ensuring stability, compliance, and risk control.

In this role you will establish SRE frameworks, define SLOs/SLIs, automate runbooks, and build auto-remediation pipelines.

Qualifications

  • Bachelor’s or Master’s degree in computer science, engineering, or a related field.
  • 5+ years in infrastructure, cloud, or SRE roles, with 3+ years in a regulated environment.
  • Hands-on expertise with observability platforms (Datadog, Dynatrace, Splunk) and IaC (Terraform, Ansible).
  • Proficient in Python and automation of runbooks and workflows.
  • Familiarity with MAS TRM, DORA, and regulatory compliance.

Responsibilities

  • Define and own enterprise observability architecture and governance.
  • Deploy and optimize full-stack observability across infra, apps, and networks.
  • Develop auto-remediation, runbooks, and anomaly detection.
  • Collaborate with Cloud, DevOps, Security to ensure compliance and audit readiness.
  • Deliver executive dashboards on availability and reliability KPIs.

Skills

Observability
SRE
Automation
Cloud
Incident governance

Education

Bachelor’s or Master’s degree in computer science, Engineering, or related discipline

Tools

Datadog
Dynatrace
Splunk
ELK
Terraform
Ansible
Python
CI/CD Tools
AWS
Azure

Job description

Associate/AVP, Observability & SRE Engineering, Technology Group

Location: Singapore, SG

Job Function: Technology Group

Job Type: Permanent

GIC is one of the world’s largest sovereign wealth funds. With over 2,000 employees across 11 locations around the world, we invest in more than 40 countries globally across asset classes and businesses. Working at GIC gives you exposure to an extraordinary network of the world’s industry leaders. As a leading global long-term investor, we work at the point of impact for Singapore’s financial future, and the communities we invest in worldwide.

Technology Group

We experiment, design, and lead a 24x7 global business where we support core capabilities in asset management, trading, investment operations, and risk management. We deliver secure, reliable, and integrated solutions, and provide insights on new and emerging technologies.

Infrastructure & Cybersecurity Resilience (ICR)

We design, build, and secure the technology foundations that power GIC’s global investment operations. We aim to deliver resilient, scalable, and secure infrastructure that empowers our people and businesses to perform securely, efficiently, and effectively.

What impact will you make in this role?

You will be responsible for executing the enterprise observability and service reliability strategy across all infrastructure and application domains within the bank. You will help ensure proactive monitoring, automation, and resilience initiatives using Datadog, Dynatrace, AWS, and Azure platforms—ensuring operational stability, risk control, and compliance with financial regulatory standards. The role also establishes SRE frameworks, Site Reliability metrics (SLO, SLI, Error Budgets), and automation pipelines to improve system reliability, reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR), and strengthen overall operational resilience.

What will you do as an Observability & SRE Engineer?
Observability Strategy & Governance
  • Define and own the Enterprise Observability Architecture aligned with operational resilience mandates (e.g., MAS TRM, DORA, APRA CPS 230).
  • Deploy and optimize observability platforms such as Datadog, Dynatrace, Splunk for full-stack visibility (infra, application, network, and user experience).
  • Establish governance standards for telemetry data—metrics, logs, and traces—ensuring consistency, retention compliance, and security controls.
  • Integrate observability platforms with incident management, ITSM, and AIOps systems for predictive alerting and anomaly detection.
Reliability Engineering & Automation
  • Contribute in implementing SRE frameworks for infrastructure and business-critical applications.
  • Drive initiatives to automate runbooks, alerts, self-healing actions, and auto-remediation workflows via Python, Ansible, and Terraform.
  • Partner with Application, Infrastructure, and Cyber teams to codify operational reliability into the delivery lifecycle.
  • Conduct resilience testing, chaos engineering, and capacity validation in alignment with business continuity standards.
  • Develop error budget policies and reliability scorecards for key production services.
Cloud Observability & Platform Engineering
  • Architect and manage observability for cloud-native workloads hosted in AWS and Azure, ensuring visibility of compute, storage, and network layers.
  • Integrate cloud observability into landing zones and CI/CD pipelines, ensuring continuous compliance with deployment controls.
  • Implement infrastructure-as-code (IaC) models using Terraform and Ansible for consistent, auditable provisioning.
  • Collaborate with Cloud, DevOps, and Security teams to ensure real-time telemetry aligns with audit and compliance requirements.
Operational Excellence & Stakeholder Management
  • Drive reduction in incident recurrence, MTTR, and manual intervention through observability-led automation.
  • Partner with Service Delivery, Cyber, and Application teams to achieve predictive incident prevention and root cause transparency.
  • Deliver executive dashboards that highlight availability, reliability KPIs, and operational risk indicators.
  • Act as a technical advisor to senior management during major incidents, post-incident reviews, and technology audits.
What makes you a successful candidate?
  • Bachelor’s or Master’s degree in computer science, Engineering, or related discipline.
  • 5+ years’ experience in Infrastructure, Cloud, or SRE roles, with 3+ years in SRE SME capacity in financial institutions or regulated environments.
  • Proven hands‑on expertise in Observability Platforms: Datadog, Dynatrace, Splunk, ELK; Automation / IaC: Terraform, Ansible, Python, CI/CD tools; Cloud Platforms: AWS (CloudWatch, X‑Ray, CloudTrail), Azure (Monitor, Log Analytics, App Insights).
  • Deep understanding of SRE principles, service health modelling, error budgets, and auto‑remediation design.
  • Familiarity with financial sector operational resilience frameworks, regulatory compliance, and incident governance.
  • Certificates in at least one of the following areas: Datadog Certified Observability Professional, Dynatrace Certified Associate, Terraform/Ansible/Python Certified Expert.
  • Highly desirable certificates: AWS Certified DevOps Engineer, Azure DevOps Expert, SRE Foundation/Practitioner (DevOps Institute), ITIL v4 Managing Professional.
GIC is an equal opportunity employer

As an employer, we passionately believe every individual brings with them unique diversity of thought and perspectives to meaningfully enrich perspectives of GIC teams to drive competitive performance. An inclusive environment yields exceptional contribution.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Associate/AVP, Observability & SRE Engineering, Technology Group
Associate/AVP, Observability & SRE Engineering, Technology Group

GIC • Singapore

Hybrid
SGD 150,000 - 195,000
AVP/VP, SIEM & SRE Engineering, Technology Group
AVP/VP, SIEM & SRE Engineering, Technology Group

GIC • Singapore

Hybrid
SGD 120,000 - 160,000
SVP, Head of IT Service Operations (Global Infrastructure & Cyber), Technology Group
SVP, Head of IT Service Operations (Global Infrastructure & Cyber), Technology Group

GIC Private Limited • Singapore

On-site
SGD 150,000 - 200,000
AVP/VP, Network Automation and Reliability Engineer, Technology Group
AVP/VP, Network Automation and Reliability Engineer, Technology Group

GIC Private Limited • Singapore

Hybrid
SGD 180,000 - 280,000
SVP, Data Platform Services Engineering Lead, Technology Group
SVP, Data Platform Services Engineering Lead, Technology Group

GIC Private Limited • Singapore

On-site
SGD 150,000 - 180,000
Flexible work arrangements
Professional growth opportunities
SVP, Head of IT Service Operations (Global Infrastructure & Cyber), Technology Group
SVP, Head of IT Service Operations (Global Infrastructure & Cyber), Technology Group

Gic • Singapore

Hybrid
SGD 150,000 - 250,000
Flexible work arrangement
Professional growth opportunities
SVP, Production Support Lead, Technology Group
SVP, Production Support Lead, Technology Group

GIC Private Limited • Singapore

On-site
SGD 120,000 - 160,000
VP/SVP, Production Support Lead (SAT), Technology Group
VP/SVP, Production Support Lead (SAT), Technology Group

GIC Private Limited • Singapore

On-site
SGD 180,000 - 240,000
AVP/VP, Incident and Business Continuity Management, Technology Group
AVP/VP, Incident and Business Continuity Management, Technology Group

GIC • Singapore

On-site
SGD 180,000 - 240,000
AVP/VP, Incident and Problem Management, Technology Group
AVP/VP, Incident and Problem Management, Technology Group

GIC Private Limited • Singapore

Hybrid
SGD 120,000 - 180,000