SA/AVP/VP - Site Reliability Engineer

SGX Group

Singapore

On-site

SGD 90,000 - 130,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

SGX Group is hiring Site Reliability Engineers who treat operations as a software problem. You will keep production healthy and build automation, tooling, and agentic workflows to make running our systems boring and predictable.

This is an engineering role in a high-reliability, regulated capital-markets environment. You will help establish SRE as a discipline, shift reliability decisions upstream, and reduce manual toil while collaborating across teams to improve resilience and incident

Qualifications

  • 5+ years in SRE, platform, or infra engineering with a track record of replacing manual work with code.
  • Strong programming ability in at least one modern language; production-ready coding.
  • AI-native ops: experience orchestrating agents for workflows, not just autocomplete.
  • Deep hands-on with Kubernetes, IaC (Terraform), CI/CD, and observability.
  • Production cloud experience (GCP preferred, AWS acceptable).
  • Foundations in distributed systems and incident-response maturity in regulated environments.

Responsibilities

  • Own production reliability (SLOs, capacity, incident response, postmortems) and turn incidents into durable fixes in code or automation.
  • Build platform and tooling for easy deployment, observability, and operation: CI/CD, IaC, observability stacks, runbooks-as-code.
  • Apply AI agentically across operations (triage, root-cause analysis, remediation, change review).
  • Design and integrate messaging, orchestration, and performance-sensitive infrastructure (e.g., Kafka, Kubernetes).
  • Partner with product engineers on release readiness, rollout strategy, and production hardening.
  • Continuously reduce toil: measure it, attack it with code, raise the floor on maintainability.

Skills

SRE experience
Go
Python
Kotlin
TypeScript
Rust
Kubernetes
Terraform
CI/CD
Observability
GCP
AWS
Distributed systems
Incident response
Regulated environments

Tools

Terraform
Kubernetes
CI/CD tooling
Observability stacks
Kafka

Job description

At SGX Group, we create markets. We turn ideas into products and expand access to new opportunities. We are building the future of the exchange in-house: from architecture and platforms to the critical systems that power markets. The biggest decisions are still open to people who join us now.

BUILD MARKETS. SHAPE ECONOMIES.

At SGX Group, we create markets. We turn ideas into products and expand access to new opportunities. We are building the future of the exchange in-house: from architecture and platforms to the critical systems that power markets. The biggest decisions are still open to people who join us now.

The opportunity

SGX is hiring Site Reliability Engineers who treat operations as a software problem. You'll keep production healthy, but more importantly you'll build the automation, tooling, and agentic workflows that make running our systems boring and predictable. This is an engineering role - if your instinct on a recurring issue is to write code that removes it, you'll fit in well.

We operate in a regulated capital-markets environment, so the bar for reliability, security, and operational rigour is high.

The Site Reliability Engineering team

Site Reliability Engineering keeps the platforms behind SGX Group's business and market infrastructure available, observable, and recoverable.

The team sets the standards other engineering teams work to across service levels, error budgets, observability, incident response and automation. It works across engineering, infrastructure, security, and product, and it owns the practice as well as seen as the SME for ensuring resilience and proactive maintenance and continuous improvement of the estate.

The next phase is about establishing SRE as a discipline rather than a function, moving reliability decisions upstream into design, and reducing the manual work that currently sits behind keeping services up.

What you will do
  • Own production reliability (SLOs, capacity, incident response, postmortems) and turn every incident into a durable fix in code or automation
  • Build the platform and tooling that make services easy to deploy, observe, and operate: CI/CD, infrastructure-as-code, observability stacks, runbooks-as-code
  • Apply AI agentically across operations (triage, root-cause analysis, remediation, change review) and contribute to our internal agentic ecosystem
  • Design and integrate the systems underneath our services: messaging (e.g. Kafka), orchestration (e.g. Kubernetes), and performance-sensitive infrastructure
  • Partner with product engineers on release readiness, rollout strategy, and production hardening before things ship
  • Continuously reduce toil: measure it, attack it with code, and raise the floor on what "easy to maintain" means here
What we are looking for
Essentials
  • 5+ years in SRE, platform, or infrastructure engineering, with a clear track record of replacing manual work with code
  • Strong programming ability in at least one modern language (e.g. Go, Python, Kotlin, TypeScript, Rust, etc), you write production code, not just glue scripts
  • AI-native ways of working: real experience orchestrating agents for ops workflows, not just using AI for autocomplete
  • Deep hands-on with Kubernetes, IaC (Terraform or equivalent), CI/CD, and modern observability (metrics, logs, traces)
  • Production experience on a major cloud: GCP preferred, AWS acceptable
  • Solid foundations in distributed systems and the failure modes that matter in production
  • Incident-response maturity: calm under pressure, sharp on root cause, disciplined about follow-through
  • Comfort in complex, regulated environments
What may set you apart
  • Familiarity with the FIX protocol or capital-markets domain
  • Experience building internal developer platforms or self-service tooling consumed by other engineers
Why this role matters

You will work on technology that underpins critical market infrastructure, where reliability is not a quality attribute of the product. It is the product. When these platforms work, participants trade and capital moves. When they do not, everyone knows within seconds.

The mandate is real. You will decide what reliability means here, how it is measured, and what the organisation is willing to trade for it. The work is demanding and the practice is still being built, which is exactly where the opportunity sits. There are not many chances in a career to establish a discipline rather than inherit one.

About SGX Group

SGX Group is one of the world's most trusted international marketplaces, known for its stability and openness. Anchored in Singapore, we enable price discovery, capital formation and risk management across asset classes, supported by resilient infrastructure and robust clearing. We convene issuers, investors and intermediaries to create and grow markets that stand the test of time. Find out more at www.SGXGroup.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

SINGAPORE EXCHANGE LIMITED • Singapore

On-site
SGD 120,000 - 160,000
SA/AVP/VP - Site Reliability Engineer
SA/AVP/VP - Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 90,000 - 130,000
AVP/VP - Site Reliability Engineer
AVP/VP - Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 120,000 - 160,000
SVP - Site Reliability Engineer
SVP - Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 300,000 - 600,000
Senior Site Reliability Engineer: AI-Driven Platform Automation
Senior Site Reliability Engineer: AI-Driven Platform Automation

SGX Group • Singapore

On-site
SGD 90,000 - 130,000
VP, Site Reliability Engineering — AI-Driven Platform & Reliability
VP, Site Reliability Engineering — AI-Driven Platform & Reliability

Singapore Exchange Limited • Singapore

On-site
SGD 120,000 - 160,000
AVP/VP - Technical Product Owner
AVP/VP - Technical Product Owner

SGX Group • Singapore

On-site
SGD 120,000 - 180,000
Asset & Wealth Management, Senior Site Reliability Engineer, Vice President, Singapore
Asset & Wealth Management, Senior Site Reliability Engineer, Vice President, Singapore

Goldman Sachs • Singapore

On-site
SGD 150,000 - 210,000
SVP - Lead Technical Product Owner (Corporate/Data Domain)
SVP - Lead Technical Product Owner (Corporate/Data Domain)

Singapore Exchange Limited • Singapore

On-site
SGD 180,000 - 240,000
Global Banking & Markets, Site Reliability Engineer, Executive Director, Singapore
Global Banking & Markets, Site Reliability Engineer, Executive Director, Singapore

goldman sachs services (singapore) pte. ltd. • Singapore

On-site
SGD 180,000 - 300,000