Get more replies from employers
Send a job-specific resume in minutes.
OCBC Group in Singapore is seeking an Observability Platform Engineer for the Site Reliability Engineering team. You will design, build, and maintain end-to-end observability across Kubernetes and OpenShift, instrument with OpenTelemetry, and implement dashboards, alerts, and SLI/SLOs to uphold service reliability in a regulated banking environment.
You will develop internal tooling in Python, create RESTful API integrations, automate with Ansible, and collaborate with DevOps and application
As Singapore's longest established bank, we have been dedicated to enabling individuals and businesses to achieve their aspirations since 1932. How? By taking the time to truly understand people. From there, we provide support, services, solutions, and career paths that meet their individual needs and desires. Today, we're on a journey of transformation. Leveraging technology and creativity to become a future-ready learning organisation. But for all that change, our strategic ambition is consistently clear and bold, which is to be Asia's leading financial services partner for a sustainable future. We invite you to build the bank of the future. Innovate the way we deliver financial services. Work in friendly, supportive teams. Build lasting value in your community. Help people grow their assets, business, and investments. Take your learning as far as you can. Or simply enjoy a vibrant, future-ready career. Your Opportunity Starts Here.
Imagine being part of a team that powers the technology behind one of Singapore's longest established banks. As a Technology Infrastructure Specialist at OCBC, you'll play a critical role in ensuring our systems and infrastructure are secure, efficient, and always available. You'll be part of a team that's driving innovation and transformation in the banking industry.
We are looking for a skilled Observability Engineer to join our Site Reliability Engineering team. In this role, you will be responsible for designing, building, and maintaining the observability platform that underpins monitoring, alerting, and operational intelligence across our critical application and infrastructure estate. You will work at the intersection of software engineering and platform operations, building tooling and automation that enables engineering teams to gain deep visibility into distributed systems running on Kubernetes and OpenShift.
Design, implement, and maintain end-to-end observability solutions covering metrics, logs, and distributed traces across Kubernetes and OpenShift clusters. Instrument applications and infrastructure using OpenTelemetry (OTel) as the standard observability framework; drive adoption across engineering teams. Build and maintain dashboards, alerting rules, and SLI/SLO frameworks to support service reliability objectives. Manage and optimise observability data pipelines, ensuring scalability, performance, and data fidelity.
Develop internal observability tooling, automation utilities, and integrations using Python. Build and maintain RESTful API integrations connecting observability platforms with internal systems and third-party services. Design and implement event-driven workflows using AWS EventBridge to enable real-time alerting, automated remediation, and cross-system notification flows. Contribute to the development of SRE Hub or equivalent service portfolio and onboarding platforms.
Develop and maintain Ansible playbooks for configuration management, observability agent deployment, and infrastructure provisioning. Automate repetitive operational tasks to reduce toil and improve platform consistency across environments. Build self-healing and auto-remediation workflows integrated into the monitoring pipeline.
Monitor and manage workloads deployed on Kubernetes and Red Hat OpenShift clusters (SIT, UAT, PROD, and DR). Configure and manage namespace-level observability, including resource monitoring, pod health, and container-level tracing. Support cluster onboarding processes, ensuring observability coverage is enforced as part of the deployment pipeline.
Work closely with DevOps and application engineering teams to embed observability as a standard within the CI/CD pipeline. Maintain codebases and automation scripts in Bitbucket; manage work items and sprint delivery through Jira. Integrate observability quality gates and checks into Jenkins pipeline stages. Participate in incident response, post-incident reviews, and the continuous improvement of operational runbooks.
The SRE team operates within a regulated banking environment. Candidates must be comfortable working with compliance, audit, and governance frameworks as part of their day-to-day responsibilities. A high-impact role within a mature SRE function supporting critical financial services infrastructure Exposure to large-scale distributed systems and enterprise observability platforms Collaborative engineering culture with a focus on automation, reliability, and continuous improvement Opportunities for growth into senior SRE, platform architecture, or engineering tracks.
Your Opportunity Starts Here