Production Support / Site Reliability Engineer- Cloud Native Data Analytics Platform

SCIENTE

Singapore

On-site

SGD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

SCIENTE in Singapore is seeking a proactive Production Support / Site Reliability Engineer to join the team supporting the Data Analytics and Reporting Platform. This role combines production support, cloud technology, data analytics, observability, automation, and financial markets knowledge.

You will own incident management, implement automated root-cause analysis, and collaborate with Trading, Risk, Infrastructure teams to improve platform reliability during market hours.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Mathematics, Finance, Engineering, or a related discipline.
  • 3-4 years of experience in supporting GCP and Google Cloud technologies (Terraform, BigQuery, Cloud Composer, Dataflow).
  • Understanding of financial markets products (Rates, Fixed Income, Repos, FX, Money Markets).
  • Experience with SRE, observability, and operational resilience.
  • Experience applying AI/ML to IT operations (AIOps, intelligent alerting, anomaly detection).
  • Production support, incident management, problem management, and operational processes.
  • Hands-on with Grafana, Prometheus, ELK/Kibana, or similar.
  • Understanding CI/CD, DevOps/SRE, and iterative software delivery.
  • Familiarity with ITIL processes in production environments.
  • Knowledge of IT risk and security concepts (SOx, certificates, authentication).
  • Strong troubleshooting and stakeholder communication skills.
  • Experience supporting business-critical platforms during market hours.

Responsibilities

  • Provide operational support for the Data Analytics and Reporting platform during Asia trading hours.
  • Own production incidents and coordinate resolutions with stakeholders.
  • Drive problem management and preventive improvements with the squad backlog.
  • Enhance platform reliability, performance, security, and observability.
  • Drive automation and AI-assisted operational improvements.
  • Support testing, release, and deployment activities for safe production changes.
  • Maintain runbooks, knowledge articles, and user guidance for onboarding.
  • Understand the full stack and its role in the platform ecosystem.
  • Collaborate with Trading, Risk, and Infrastructure teams to improve stability.
  • Participate in a global standby rotation and Agile/Scrum backlog reviews.

Skills

GCP
Terraform
BigQuery
Cloud Composer
Dataflow
Observability
SRE principles
Incident management
Prometheus
Grafana
ELK/Kibana
CI/CD
ITIL awareness
Security concepts

Education

Bachelor's or Master's degree in Computer Science/Math/Engineering

Tools

Grafana
Prometheus
ELK/Kibana

Job description

We are looking for a proactive and technically strong Production Support / Site Reliability Engineer to join the team supporting the Data Analytics and Reporting Platform. This is a business-critical role combining production support, cloud technology, data analytics, observability, automation, and financial markets knowledge.


Mandatory Skill-set


  • Bachelor's or Master's degree in Computer Science, Mathematics, Finance, Engineering, or a related discipline;

  • Must have 3-4 years of experience in supporting , GCP, Google Cloud technologies such as Terraform, BigQuery, Cloud Composer, Dataflow and cloud-native data analytics platforms;

  • Understanding of financial markets products, particularly, Rates & Credits, Fixed Income, Repos, Foreign Exchange (FX), Money Markets;

  • Experience with SRE principles, service reliability engineering, observability, and operational resilience;

  • Experience applying AI or machine learning to IT operations, such as AIOps, intelligent alerting, incident summarization, anomaly detection, or automated root-cause analysis

  • Experience with production support, incident management, problem management, and operational processes;

  • Hands-on experience with monitoring and observability tools such as Grafana, Prometheus, ELK/Kibana, or equivalent technologies;

  • Understanding of CI/CD, automation, DevOps/SRE practices, and iterative software delivery;

  • Understanding of ITIL processes and their practical application in a production environment;

  • Knowledge of IT risk and security concepts, including SOx controls, vulnerability management, certificates, authentication, and secure communication protocols;

  • Strong troubleshooting and analytical skills, with the ability to understand complex systems and identify root causes;

  • Experience supporting business-critical platforms during market/trading hours;

  • Strong communication and stakeholder-management skills, with the ability to communicate effectively with both technical and non-technical audiences.


Desired Skill-set


  • Experience applying AI or machine learning to IT operations, such as AIOps, intelligent alerting, incident summarization, anomaly detection, or automated root-cause analysis;

  • ITIL certification.


Responsibilities


  • Provide operational support for the Data Analytics and Reporting platform, ensuring reliable service delivery during Asia trading hours;

  • Take ownership of production incidents, coordinating resolution activities and stakeholder communication to minimize business impact;

  • Drive effective problem management by identifying recurring issues, implementing preventive measures, and collaborating with the Product Owner to prioritize improvements through the squad backlog;

  • Contribute to the continuous enhancement of the platform with a strong focus on reliability, performance, security, and operational excellence;

  • Drive automation and AI-assisted operational improvements to enhance platform reliability, efficiency, and observability;

  • Support testing, release, and deployment activities to ensure safe and stable delivery of changes into production;

  • Maintain and continuously improve operational runbooks, knowledge articles, and user guidance documentation to support platform onboarding;

  • You understand the entire stack’s technology on which the application runs and how it fits in the overall chain;

  • Collaborate closely with Trading, Risk, Infrastructure teams and other squads to resolve complex issues and continuously improve platform stability and user experience;

  • Participate in a shared global standby rotation (typically one week per month), contributing to the reliability and continuity of business-critical services;

  • As part of a cross-border Squad work in an Agile/Scrum way, on the backlog prioritized by a Product Owner, and demonstrate your features/stories to other colleagues and the stakeholders.


Confidentiality is assured, and only shortlisted candidates will be notified for interviews.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Support / Site Reliability Engineer- Cloud Native Data Analytics Platform (JD#11280)
Production Support / Site Reliability Engineer- Cloud Native Data Analytics Platform (JD#11280)

SCIENTE INTERNATIONAL PTE. LTD. • Singapore

On-site
SGD 120,000 - 190,000
Operations Support Engineer - GCP
Operations Support Engineer - GCP

SCIENTE • Singapore

On-site
SGD 90,000 - 130,000
Operations Support Engineer - GCP (JD#11278)
Operations Support Engineer - GCP (JD#11278)

SCIENTE INTERNATIONAL PTE. LTD. • Singapore

On-site
SGD 70,000 - 110,000
Application Support Engineer (Google Cloud Platform/GCP, Financial Markets)
Application Support Engineer (Google Cloud Platform/GCP, Financial Markets)

BEATHCHAPMAN (PTE. LTD.) • Singapore

On-site
SGD 90,000 - 130,000
SRE & Production Support Engineer - Cloud & Observability
SRE & Production Support Engineer - Cloud & Observability

SCIENTE INTERNATIONAL PTE. LTD. • Singapore

On-site
SGD 120,000 - 190,000
Technology Support Lead, Prime Finance and Clearing Production Management Support
Technology Support Lead, Prime Finance and Clearing Production Management Support

JPMorgan Chase & Co. • Singapore

On-site
SGD 80,000 - 120,000
Senior Site Reliability Engineer - Data Platform & AI Ops
Senior Site Reliability Engineer - Data Platform & AI Ops

SCIENTE • Singapore

On-site
SGD 120,000 - 180,000
GCP Operations & Data Analytics Engineer
GCP Operations & Data Analytics Engineer

SCIENTE INTERNATIONAL PTE. LTD. • Singapore

On-site
SGD 70,000 - 110,000
Front Office Equity Support
Front Office Equity Support

Smart IMS Inc. • Singapore

On-site
SGD 150,000 - 190,000
Production Specialist - Digital FX
Production Specialist - Digital FX

NATWEST MARKETS PLC • Singapore

On-site
SGD 120,000 - 190,000