AVP/VP - Site Reliability Engineer

United States Digital Space LLC

Singapore

On-site

SGD 120,000 - 180,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Singapore Exchange is seeking Site Reliability Engineers to treat operations as a software problem and keep production healthy through automation, tooling, and agentic workflows. This is an engineering role in a regulated capital-markets environment with high reliability and security bar.

You will own production reliability, build platforms for easy deployment, and apply AI-native approaches across operations, including triage, root-cause analysis, and remediation.

Qualifications

  • 5+ years in SRE, platform, or infrastructure engineering.
  • Strong programming ability in at least one modern language.
  • AI-native approaches to operations and agentic workflows.
  • Production cloud experience with distributed systems and resilience.

Responsibilities

  • Own production reliability (SLOs, capacity, incident response, postmortems) and turn every incident into a durable fix in code or automation.
  • Build the platform and tooling for easy deployment, observability, and operation: CI/CD, infrastructure-as-code, observability stacks, runbooks-as-code.
  • Apply AI-agentic approaches across operations (triage, root-cause analysis, remediation, change review).
  • Design and integrate systems underneath services: messaging (Kafka), orchestration (Kubernetes), and scalable infrastructure.
  • Partner with product engineers on release readiness, rollout strategy, and production hardening before things ship.
  • Continuously reduce toil: measure it, attack it with code, raise the floor on ease of maintenance.

Skills

Go
Python
Kotlin
TypeScript
Rust

Tools

Kubernetes
Terraform
CI/CD
Observability
Kafka

Job description

Company: Singapore Exchange

Location: Singapore, SG

Unit: Technology

Job Type: Permanent (HC)

Requisition ID: 3475

Job Summary

the company is hiring Site Reliability Engineers who treat operations as a software problem. You'll keep production healthy, but more importantly you'll build the automation, tooling, and agentic workflows that make running our systems boring and predictable. This is an engineering role - if your instinct on a recurring issue is to write code that removes it, you'll fit in well.

We operate in a regulated capital-markets environment, so the bar for reliability, security, and operational rigour is high.

Job Responsibilities
  • Own production reliability (SLOs, capacity, incident response, postmortems) and turn every incident into a durable fix in code or automation
  • Build the platform and tooling that make services easy to deploy, observe, and operate: CI/CD, infrastructure-as-code, observability stacks, runbooks-as-code
  • Apply AI agentically across operations (triage, root-cause analysis, remediation, change review) and contribute to our internal agentic ecosystem
  • Design and integrate the systems underneath our services: messaging (e.g. Kafka), orchestration (e.g. Kubernetes), and performance-sensitive infrastructure
  • Partner with product engineers on release readiness, rollout strategy, and production hardening before things ship
  • Continuously reduce toil: measure it, attack it with code, and raise the floor on what "easy to maintain" means here
Job Requirements
  • 5+ years in SRE, platform, or infrastructure engineering, with a clear track record of replacing manual work with code
  • Strong programming ability in at least one modern language (e.g. Go, Python, Kotlin, TypeScript, Rust, etc), you write production code, not just glue scripts
  • AI-native ways of working: real experience orchestrating agents for ops workflows, not just using AI for autocomplete
  • Deep hands-on with Kubernetes, IaC (Terraform or equivalent), CI/CD, and modern observability (metrics, logs, traces)
  • Production experience on a major cloud: GCP preferred, AWS acceptable
  • Solid foundations in distributed systems and the failure modes that matter in production
  • Incident-response maturity: calm under pressure, sharp on root cause, disciplined about follow-through
  • Comfort in complex, regulated environments
Nice to Have

Familiarity with the FIX protocol or capital-markets domainExperience building internal developer platforms or self-service tooling consumed by other engineers

Job Segment: Cloud, Executive, VP, Technology, Management

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AVP/VP - Site Reliability Engineer
AVP/VP - Site Reliability Engineer

SGX Group • Singapore

On-site
SGD 120,000 - 170,000
AVP/VP - Site Reliability Engineer
AVP/VP - Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

SGX Group • Singapore

On-site
SGD 180,000 - 300,000
SA/AVP/VP - Site Reliability Engineer
SA/AVP/VP - Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 90,000 - 130,000
VP, Site Reliability Engineering — AI-Driven Platform & Reliability
VP, Site Reliability Engineering — AI-Driven Platform & Reliability

Singapore Exchange Limited • Singapore

On-site
SGD 120,000 - 160,000
SVP, Site Reliability Engineer
SVP, Site Reliability Engineer

SGX Group • Singapore

On-site
SGD 300,000 - 420,000
SVP, Site Reliability Engineer
SVP, Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 350,000 - 700,000
VP/Senior SRE – Platform & Automation
VP/Senior SRE – Platform & Automation

SGX Group • Singapore

On-site
SGD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 180,000 - 250,000
AVP Site Reliability Platform Lead
AVP Site Reliability Platform Lead

Singapore Exchange Limited • Singapore

On-site
SGD 140,000 - 210,000