Observability Engineer - Site Reliability Engineering (SRE)

United States Digital Space LLC

Singapore

On-site

SGD 90,000 - 120,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive base salary
Holistic, flexible benefits
Industry‑leading learning and dev

Job summary

Singapore’s longest established bank is seeking an Observability Engineer to join the Site Reliability Engineering team. You will design, build, and maintain the observability platform across Kubernetes and OpenShift to ensure systems are secure, efficient, and highly available.

Key tasks include instrumenting apps with OpenTelemetry, dashboards and SLI/SLOs, and developing internal tooling in Python. You’ll integrate with AWS EventBridge, RESTful APIs, and CI/CD pipelines, partnering with

Qualifications

  • 3–6 years of experience in an SRE, Platform Engineering, or DevOps role with a strong observability focus.
  • Hands‑on experience with Kubernetes and/or Red Hat OpenShift — cluster operations, namespaces, workloads, and resource management
  • Proficient in Python — production‑quality scripts, automation utilities, and API integrations
  • Working experience with OpenTelemetry (OTel) — instrumentation, collectors, exporters, and signal types (logs, metrics, traces)
  • Experience building or consuming RESTful APIs
  • Familiarity with AWS EventBridge for event‑driven architecture and integration workflows
  • Experience with Ansible for configuration management and infrastructure automation
  • Working knowledge of DevOps tooling: Bitbucket (Git), Jira, Jenkins

Responsibilities

  • Design, implement, and maintain end‑to‑end observability solutions covering metrics, logs, and distributed traces across Kubernetes and OpenShift clusters.
  • Instrument applications and infrastructure using OpenTelemetry (OTel) as the standard observability framework; drive adoption across engineering teams.
  • Build and maintain dashboards, alerting rules, and SLI/SLO frameworks to support service reliability objectives.
  • Manage and optimise observability data pipelines, ensuring scalability, performance, and data fidelity.
  • Develop internal observability tooling, automation utilities, and integrations using Python.
  • Build and maintain RESTful API integrations connecting observability platforms with internal systems and third‑party services.
  • Design and implement event‑driven workflows using AWS EventBridge to enable real‑time alerting, automated remediation, and cross‑system notification flows.
  • Contribute to the development of SRE Hub or equivalent service portfolio and onboarding platforms.
  • Monitor and manage workloads deployed on Kubernetes and Red Hat OpenShift clusters, including namespace observability and container tracing.

Skills

SRE / Platform Engineering
DevOps
Python
OpenTelemetry
Kubernetes
OpenShift
RESTful APIs
AWS EventBridge
Ansible
Bitbucket
Jira
Jenkins

Education

Bachelor's degree in Computer Science / IT / Engineering

Tools

Bitbucket
Jira
Jenkins
Ansible

Job description

WHO WE ARE:

As Singapore’s longest established bank, we have been dedicated to enabling individuals and businesses to achieve their aspirations since 1932. How? By taking the time to truly understand people. From there, we provide support, services, solutions, and career paths that meet their individual needs and desires.

Today, we’re on a journey of transformation. Leveraging technology and creativity to become a future-ready learning organisation. But for all that change, our strategic ambition is consistently clear and bold, which is to be Asia’s leading financial services partner for a sustainable future.

We invite you to build the bank of the future. Innovate the way we deliver financial services. Work in friendly, supportive teams. Build lasting value in your community. Help people grow their assets, business, and investments. Take your learning as far as you can. Or simply enjoy a vibrant, future-ready career.

Your Opportunity Starts Here.

Why Join

Imagine being part of a team that powers the technology behind one of Singapore's longest established banks. As a Technology Infrastructure Specialist at the company, you'll play a critical role in ensuring our systems and infrastructure are secure, efficient, and always available. You'll be part of a team that's driving innovation and transformation in the banking industry.

How you succeed

We are looking for a skilled Observability Engineer to join our Site Reliability Engineering team. In this role, you will be responsible for designing, building, and maintaining the observability platform that underpins monitoring, alerting, and operational intelligence across our critical application and infrastructure estate. You will work at the intersection of software engineering and platform operations, building tooling and automation that enables engineering teams to gain deep visibility into distributed systems running on Kubernetes and OpenShift.

What you do
Observability Platform Engineering
  • Design, implement, and maintain end-to-end observability solutions covering metrics, logs, and distributed traces across Kubernetes and OpenShift clusters.
  • Instrument applications and infrastructure using OpenTelemetry (OTel) as the standard observability framework; drive adoption across engineering teams.
  • Build and maintain dashboards, alerting rules, and SLI/SLO frameworks to support service reliability objectives.
  • Manage and optimise observability data pipelines, ensuring scalability, performance, and data fidelity.
Application & API Development
  • Develop internal observability tooling, automation utilities, and integrations using Python.
  • Build and maintain RESTful API integrations connecting observability platforms with internal systems and third‑party services.
  • Design and implement event‑driven workflows using AWS EventBridge to enable real‑time alerting, automated remediation, and cross‑system notification flows.
  • Contribute to the development of SRE Hub or equivalent service portfolio and onboarding platforms.
Automation & Configuration Management
  • Develop and maintain Ansible playbooks for configuration management, observability agent deployment, and infrastructure provisioning.
  • Automate repetitive operational tasks to reduce toil and improve platform consistency across environments.
  • Build self‑healing and auto‑remediation workflows integrated into the monitoring pipeline.
Kubernetes & OpenShift Operations
  • Monitor and manage workloads deployed on Kubernetes and Red Hat OpenShift clusters (SIT, UAT, PROD, and DR).
  • Configure and manage namespace‑level observability, including resource monitoring, pod health, and container‑level tracing.
  • Support cluster onboarding processes, ensuring observability coverage is enforced as part of the deployment pipeline.
DevOps Collaboration
  • Work closely with DevOps and application engineering teams to embed observability as a standard within the CI/CD pipeline.
  • Maintain codebases and automation scripts in Bitbucket; manage work items and sprint delivery through Jira.
  • Integrate observability quality gates and checks into Jenkins pipeline stages.
  • Participate in incident response, post‑incident reviews, and the continuous improvement of operational runbooks.
Who you are
Qualifications
  • Bachelor's degree in Computer Science, Information Technology, or a related engineering discipline
  • Relevant certifications are advantageous: CKA/CKAD (Kubernetes), Red Hat OpenShift, AWS Solutions Architect, or equivalent
Must Have
  • 3–6 years of experience in an SRE, Platform Engineering, or DevOps role with a strong observability focus
  • Hands‑on experience with Kubernetes and/or Red Hat OpenShift — cluster operations, namespaces, workloads, and resource management
  • Proficient in Python — ability to write production‑quality scripts, automation utilities, and API integrations
  • Working experience with OpenTelemetry (OTel) — instrumentation, collectors, exporters, and signal types (logs, metrics, traces)
  • Experience building or consuming RESTful APIs
  • Familiarity with AWS services, particularly AWS EventBridge for event‑driven architecture and integration workflows
  • Experience with Ansible for configuration management and infrastructure automation
  • Working knowledge of DevOps tooling: Bitbucket (Git), Jira, Jenkins
Good to Have
  • Experience with Elastic Stack (ELK) — Elasticsearch, Logstash, Kibana, Elastic APM, Elastic Watcher
  • Familiarity with Moogsoft, Dynatrace, Prometheus, Grafana, or similar observability platforms
  • Understanding of service mesh (Istio) and its observability capabilities
  • Exposure to SLI/SLO frameworks and error budget management
  • Experience with cloud‑native CI/CD pipelines and GitOps practices
  • Knowledge of microservice architecture and distributed tracing in production environments
  • Familiarity with regulated or financial services environments
What We Offer

The SRE team operates within a regulated banking environment. Candidates must be comfortable working with compliance, audit, and governance frameworks as part of their day‑to‑day responsibilities.

  • A high-impact role within a mature SRE function supporting critical financial services infrastructure
  • Exposure to large-scale distributed systems and enterprise observability platforms
  • Collaborative engineering culture with a focus on automation, reliability, and continuous improvement
  • Opportunities for growth into senior SRE, platform architecture, or engineering tracks
Who we are

As Singapore's longest established bank, we have been dedicated to enabling individuals and businesses to achieve their aspirations since 1932. How? By taking the time to truly understand people. From there, we provide support, services, solutions, and career paths that meet their individual needs and desires.

Today, we're on a journey of transformation. Leveraging technology and creativity to become a future‑ready learning organisation.

But for all that change, our strategic ambition is consistently clear and bold, which is to be Asia's leading financial services partner for a sustainable future.

We invite you to build the bank of the future. Innovate the way we deliver financial services. Work in friendly, supportive teams. Build lasting value in your community. Help people grow their assets, business, and investments. Take your learning as far as you can. Or simply enjoy a vibrant, future‑ready career. Your Opportunity Starts Here.

What we offer:
  • Competitive base salary
  • A suite of holistic, flexible benefits to suit every lifestyle
  • Community initiatives
  • Industry‑leading learning and professional development opportunities
  • Your wellbeing, growth and aspirations are every bit as cared for as the needs of our customers
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Engineer - Site Reliability Engineering (SRE)
Observability Engineer - Site Reliability Engineering (SRE)

OCBC • Singapore

On-site
SGD 120,000 - 180,000
Competitive base salary
Holistic benefits package
Learning and development opportunities
Observability Engineer - Site Reliability Engineering (SRE)
Observability Engineer - Site Reliability Engineering (SRE)

OCBC Group • Singapore

On-site
SGD 120,000 - 180,000
Competitive base salary.
Flexible benefits package
Community initiatives.
+2
Observability Engineer - SRE, Kubernetes & Automation
Observability Engineer - SRE, Kubernetes & Automation

United States Digital Space LLC • Singapore

On-site
SGD 90,000 - 120,000
Competitive base salary
Holistic, flexible benefits
Industry‑leading learning and dev
Associate/AVP, Observability & SRE Engineering, Technology Group
Associate/AVP, Observability & SRE Engineering, Technology Group

GIC Private Limited • Singapore

On-site
SGD 150,000 - 190,000
Senior Associate, Observability Platform SRE Engineer, SRE & Governance, Group Technology
Senior Associate, Observability Platform SRE Engineer, SRE & Governance, Group Technology

DBS Bank Ltd • Singapore

On-site
SGD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

BEATHCHAPMAN (PTE. LTD.) • Singapore

Hybrid
SGD 120,000 - 180,000
Hybrid working arrangement
Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

Singapore Pools • Singapore

On-site
SGD 120,000 - 170,000
Total rewards package
Health & wellness benefits
Continuous learning and upskilling
+1
DevOps Engineer (AVP)
DevOps Engineer (AVP)

United States Digital Space LLC • Singapore

On-site
SGD 140,000 - 220,000
Competitive base salary
Flexible benefits
Learning & development
+1
Observability Engineer — SRE for Kubernetes/OpenShift
Observability Engineer — SRE for Kubernetes/OpenShift

OCBC Group • Singapore

On-site
SGD 120,000 - 180,000
Competitive base salary.
Flexible benefits package
Community initiatives.
+2
Site Reliability Engineer - Data Availability
Site Reliability Engineer - Data Availability

SIX Group Services Ltd. • Singapore

Hybrid
SGD 70,000 - 90,000
Flexible work models
Personal development opportunities
Agile working methods