Site Reliability Engineer

BEATHCHAPMAN (PTE. LTD.)

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid working arrangement

Job summary

BeathChapman Pte Ltd in Singapore is seeking an experienced Site Reliability Engineer to join the SRE team. The role starts hands-on as an individual contributor with a genuine path to grow into a lead position, reporting to the Head of SRE.

You will own the production platforms for FX and payments, driving reliability with strong observability and automation. The role offers exposure to senior engineering leadership and a modern tech environment with AI-assisted tooling, plus a hybrid working

Qualifications

  • Hands-on 5+ years in site reliability, production support or platform engineering.
  • Strong troubleshooting across production systems, logs, APIs and behavior.
  • Comfort with coding/scripting (Python, Bash, Java or similar).
  • Cloud and containers experience (AWS, Docker, Kubernetes) and monitoring tooling.

Responsibilities

  • Own the production platform day to day - investigate alerts, diagnose issues from logs, and coordinate with product, ops and engineering teams to resolve production and integration issues.
  • Own production reliability across the FX and payments platforms - monitoring, observability and alerting, including refining alerts as the platform evolves.
  • Lead incident response and root-cause analysis, escalating genuine application bugs to developers while resolving configuration and platform issues directly.
  • Support UAT and quality testing, and help strengthen test coverage and release standards toward automation and AI-assisted workflows.
  • Support client API integration and onboarding - guiding clients through the platform and acting as the technical bridge between clients, ops and engineering.
  • Contribute to DevOps, tooling and monitoring improvements across reliability, testing and support.
  • Apply incident management, SLA and BCP/DR practice, maintaining runbooks and operational discipline.
  • As the role grows, coach and uplift junior SRE and QA engineers and help set standards as the function matures.

Skills

Python
Troubleshooting
Incident management
Collaboration
Client-facing

Tools

AWS
Docker
Kubernetes
Grafana
Prometheus
OpenSearch
CloudWatch

Job description

Client introduction

Our client is an established fintech headquartered in Singapore, operating across payments and foreign exchange with a footprint spanning Asia and beyond, including a significant client base in Greater China. As the business scales towards its next stage of growth, they are building out the reliability function that keeps their production platforms running.

This is a newly created Site Reliability Engineer role within the SRE team. It starts hands‑on as an individual contributor, reporting to the Head of SRE, with a genuine path to grow into a lead position over time as part of succession planning.

Job responsibilities
  • Own the production platform day to day - investigate alerts, diagnose issues from logs, make technical configuration changes where needed, and coordinate with product, ops and engineering teams to resolve production and integration issues.
  • Own production reliability across the FX and payments platforms - monitoring, observability and alerting, including refining alerts as the platform evolves.
  • Lead incident response and root-cause analysis, escalating genuine application bugs to developers while resolving configuration and platform issues directly.
  • Support UAT and quality testing alongside the team, and help strengthen test coverage and release standards, moving toward more automated and AI-assisted workflows.
  • Support client API integration and onboarding - stepping in on technical integration issues, guiding clients through the platform, and acting as the technical bridge between clients, ops and engineering.
  • Contribute to DevOps, tooling and monitoring improvements across reliability, testing and support.
  • Apply sound incident management, SLA and BCP/DR practice, maintaining runbooks and operational discipline.
  • As the role grows, coach and uplift junior SRE and QA engineers and help set standards as the function matures.
Job requirements
  • At least 5 years of hands‑on experience in site reliability, production support (L2), platform engineering or technical operations, ideally within fintech, payments, FX or another high‑availability environment.
  • Strong hands‑on troubleshooting across production systems, logs, APIs and application behaviour - able to diagnose issues and make technical configuration changes independently.
  • Comfortable coding and scripting (Python, Bash, Java or similar) - hands‑on coding is part of the role.
  • Solid fundamentals across cloud and containers (AWS, Docker, Kubernetes) and monitoring/observability tooling (Grafana, Prometheus, OpenSearch, CloudWatch).
  • Sound grounding in incident management, RCA, SLA and BCP/DR practice.
  • Quality assurance exposure (test automation, release or regression testing) is an advantage - the role includes a QA element alongside its reliability core.
  • Working knowledge of API integration and comfort in client-facing or client-support situations is a plus; the role includes some direct client interaction.
  • Someone hands‑on today, keen to learn and take the next step, who wants to grow into a broader leadership role. People‑management or mentoring experience is an advantage but not required - the role starts as an individual contributor and grows into leading a small team.
  • Candidates must be legally eligible to work in Singapore. Relocation support may be considered for suitable candidates based overseas.
Why you should join them
  • A newly created role with a genuine path to grow into a lead position, as part of real succession planning within the reliability function.
  • Ownership of the production platform from day one - you'll be the person the alerts come to, with room to shape how reliability, quality and support run rather than inherit a fixed playbook.
  • Direct exposure to senior engineering leadership, reporting to the Head of SRE and working closely with the wider engineering and infrastructure teams.
  • A modern technical environment with real appetite for AI‑assisted tooling across testing, RCA and support, plus a hybrid working arrangement.

JL
Reg No. R1766249
BeathChapman Pte Ltd
Licence no. 16S8112

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TEKsystems • Singapore

Hybrid
SGD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Dexian • Singapore

Hybrid
SGD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Allegis Group Singapore Pte Ltd • Singapore

Hybrid
SGD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

ALLEGIS GROUP SINGAPORE PRIVATE LIMITED • Singapore

Hybrid
SGD 120,000 - 180,000
Site Reliability Engineer – AI‑Driven, Growth Path
Site Reliability Engineer – AI‑Driven, Growth Path

BEATHCHAPMAN (PTE. LTD.) • Singapore

Hybrid
SGD 120,000 - 180,000
Hybrid working arrangement
Site Reliability Engineer - Data Availability
Site Reliability Engineer - Data Availability

SIX Group Services Ltd. • Singapore

Hybrid
SGD 70,000 - 90,000
Flexible work models
Personal development opportunities
Agile working methods
AVP/VP - Site Reliability Engineer
AVP/VP - Site Reliability Engineer

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

SGX Group • Singapore

On-site
SGD 180,000 - 300,000
Site Reliability Engineer
Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 180,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

Auxo Talent • Singapore

On-site
SGD 90,000 - 150,000