Senior Site Reliability Engineer (SRE)

GFT Technologies APAC & GCC

Hong Kong

On-site

HKD 900,000 - 1,500,000

Full time

13 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Professional & fun working environment
Growth opportunities
Future-ready digital bank platform

Job summary

GFT Technologies seeks a Senior Site Reliability Engineer to support a Stablecoin and Digital Assets Platform. The role emphasizes secure, reliable production operations, observability, and incident response across cloud-native environments.

You will partner with Product, Engineering, Infrastructure, Security, Compliance, Risk, and Operations to improve resilience and operational maturity while supporting 24x7 production services.

Qualifications

  • 7+ years in SRE, production support, or related roles in fintech or regulated environments.
  • Strong skills in monitoring, observability, alerting, and incident management.
  • Experience with cloud platforms, preferably AWS.
  • Familiarity with CI/CD pipelines and infrastructure automation.
  • Proven problem-solving abilities in financial services settings.

Responsibilities

  • Own day-to-day production support across platform environments.
  • Enhance observability, dashboards, logging, and telemetry.
  • Coordinate incident management and post-mortems.
  • Support CI/CD pipelines and infrastructure automation.
  • Collaborate with Security, Compliance, Risk, and Ops teams.

Skills

SRE
Production Support
Platform Operations
DevOps
Cloud AWS
CI/CD
Troubleshooting
Financial services

Tools

Grafana
OpenTelemetry
CloudWatch
ELK
Datadog
Splunk
Kubernetes
Linux
Python
Bash

Job description

GFT Technologies is an AI-centric global digital transformation company. We design advanced data and AI transformation solutions, modernize technology architectures and develop next-generation core systems for industry leaders in Banking, Insurance, Manufacturing and Robotics. Partnering closely with our clients, we push boundaries to unlock their full potential. With deep industry expertise, cutting-edge technology, and a strong partner ecosystem, GFT delivers responsible AI-centric solutions that combine engineering excellence, high-performance delivery, and cost efficiency. Our team of 12,000+ technology experts operate in 20+ countries worldwide offering career opportunities at the forefront of software innovation.

Overview

We are seeking an experienced Senior Site Reliability Engineer (SRE) / Production Operations Engineer to support the production operations and operational readiness of a next-generation Stablecoin and Digital Assets Platform.

This role is focused on ensuring the platform operates securely, reliably, and efficiently in production while enhancing observability, monitoring, alerting, and operational support capabilities.

The successful candidate will work closely with Product, Engineering, Infrastructure, Security, Compliance, Risk, and Operations teams to support production environments, troubleshoot issues, improve operational processes, and continuously improve platform reliability and resilience.

Key Responsibilities
Production Operations
  • Own day-to-day production support activities across platform environments.
  • Monitor platform availability, performance, and operational health.
  • Investigate and troubleshoot production issues across application, infrastructure, third-party services, and blockchain components.
  • Coordinate incident management, root cause analysis, and issue resolution.
  • Participate in after-hours production support and critical incident response when required.
Observability & Reliability
  • Design and enhance platform observability, monitoring, and alerting capabilities.
  • Implement centralized dashboards, logging, and operational visibility across platform components.
  • Improve incident detection, troubleshooting, and root cause analysis through effective use of logs, metrics, and telemetry.
  • Identify reliability risks and drive improvements in platform resilience and operational maturity.
Infrastructure & Platform Operations
  • Support cloud infrastructure and production environments.
  • Support deployment processes, CI/CD pipelines, and infrastructure automation.
  • Work closely with engineering teams to improve operational stability and release quality.
  • Assist in diagnosing infrastructure and performance-related issues.
Security, Compliance & Stakeholder Coordination
  • Support operational controls required for a regulated financial platform.
  • Participate in security reviews, audits, disaster recovery exercises, and operational readiness activities.
  • Collaborate with Engineering, Product, Operations, Compliance, Risk, and external technology providers.
  • Provide clear operational reporting and incident updates to stakeholders.
Required Experience
  • 7+ years of experience in Site Reliability Engineering (SRE), Production Support, Platform Operations, DevOps, Infrastructure Engineering, or similar roles.
  • Experience supporting business-critical production systems.
  • Strong experience with monitoring, observability, alerting, incident management, and production troubleshooting.
  • Experience with cloud platforms, preferably AWS.
  • Familiarity with CI/CD pipelines and infrastructure automation.
  • Strong troubleshooting and problem-solving skills.
  • Experience working in financial services, fintech, payments, digital banking, or other regulated environments.
Preferred Experience
  • Experience supporting blockchain, digital asset, stablecoin, payment, or fintech platforms.
  • Experience with observability tools such as Grafana, OpenTelemetry, CloudWatch, ELK, Datadog, Splunk, or similar platforms.
  • Experience with Kubernetes and containerized workloads.
  • Experience with Linux administration.
  • Experience with scripting and automation (Python, Bash, PowerShell, or similar).
  • Experience operating or supporting 24x7 production services.
Profile description
Required Experience
  • 7+ years of experience in Site Reliability Engineering (SRE), Production Support, Platform Operations, DevOps, Infrastructure Engineering, or similar roles.
  • Experience supporting business-critical production systems.
  • Strong experience with monitoring, observability, alerting, incident management, and production troubleshooting.
  • Experience with cloud platforms, preferably AWS.
  • Familiarity with CI/CD pipelines and infrastructure automation.
  • Strong troubleshooting and problem-solving skills.
  • Experience working in financial services, fintech, payments, digital banking, or other regulated environments.
Preferred Experience
  • Experience supporting blockchain, digital asset, stablecoin, payment, or fintech platforms.
  • Experience with observability tools such as Grafana, OpenTelemetry, CloudWatch, ELK, Datadog, Splunk, or similar platforms.
  • Experience with Kubernetes and containerized workloads.
  • Experience with Linux administration.
  • Experience with scripting and automation (Python, Bash, PowerShell, or similar).
  • Experience operating or supporting 24x7 production services.
Job description
About GFT

GFT Technologies is an AI-centric global digital transformation company. We design advanced data and AI transformation solutions, modernize technology architectures and develop next-generation core systems for industry leaders in Banking, Insurance, Manufacturing and Robotics. Partnering closely with our clients, we push boundaries to unlock their full potential. With deep industry expertise, cutting-edge technology, and a strong partner ecosystem, GFT delivers responsible AI-centric solutions that combine engineering excellence, high-performance delivery, and cost efficiency. Our team of 12,000+ technology experts operate in 20+ countries worldwide offering career opportunities at the forefront of software innovation.

Overview

We are seeking an experienced Senior Site Reliability Engineer (SRE) / Production Operations Engineer to support the production operations and operational readiness of a next-generation Stablecoin and Digital Assets Platform.

This role is focused on ensuring the platform operates securely, reliably, and efficiently in production while enhancing observability, monitoring, alerting, and operational support capabilities.

The successful candidate will work closely with Product, Engineering, Infrastructure, Security, Compliance, Risk, and Operations teams to support production environments, troubleshoot issues, improve operational processes, and continuously improve platform reliability and resilience.

Key Responsibilities
Production Operations
  • Own day-to-day production support activities across platform environments.
  • Monitor platform availability, performance, and operational health.
  • Investigate and troubleshoot production issues across application, infrastructure, third-party services, and blockchain components.
  • Coordinate incident management, root cause analysis, and issue resolution.
  • Participate in after-hours production support and critical incident response when required.
Observability & Reliability
  • Design and enhance platform observability, monitoring, and alerting capabilities.
  • Implement centralized dashboards, logging, and operational visibility across platform components.
  • Improve incident detection, troubleshooting, and root cause analysis through effective use of logs, metrics, and telemetry.
  • Identify reliability risks and drive improvements in platform resilience and operational maturity.
Infrastructure & Platform Operations
  • Support cloud infrastructure and production environments.
  • Support deployment processes, CI/CD pipelines, and infrastructure automation.
  • Work closely with engineering teams to improve operational stability and release quality.
  • Assist in diagnosing infrastructure and performance-related issues.
Security, Compliance & Stakeholder Coordination
  • Support operational controls required for a regulated financial platform.
  • Participate in security reviews, audits, disaster recovery exercises, and operational readiness activities.
  • Collaborate with Engineering, Product, Operations, Compliance, Risk, and external technology providers.
  • Provide clear operational reporting and incident updates to stakeholders.
Required Experience
  • 7+ years of experience in Site Reliability Engineering (SRE), Production Support, Platform Operations, DevOps, Infrastructure Engineering, or similar roles.
  • Experience supporting business-critical production systems.
  • Strong experience with monitoring, observability, alerting, incident management, and production troubleshooting.
  • Experience with cloud platforms, preferably AWS.
  • Familiarity with CI/CD pipelines and infrastructure automation.
  • Strong troubleshooting and problem-solving skills.
  • Experience working in financial services, fintech, payments, digital banking, or other regulated environments.
Preferred Experience
  • Experience supporting blockchain, digital asset, stablecoin, payment, or fintech platforms.
  • Experience with observability tools such as Grafana, OpenTelemetry, CloudWatch, ELK, Datadog, Splunk, or similar platforms.
  • Experience with Kubernetes and containerized workloads.
  • Experience with Linux administration.
  • Experience with scripting and automation (Python, Bash, PowerShell, or similar).
  • Experience operating or supporting 24x7 production services.
Profile description
Required Experience
  • 7+ years of experience in Site Reliability Engineering (SRE), Production Support, Platform Operations, DevOps, Infrastructure Engineering, or similar roles.
  • Experience supporting business-critical production systems.
  • Strong experience with monitoring, observability, alerting, incident management, and production troubleshooting.
  • Experience with cloud platforms, preferably AWS.
  • Familiarity with CI/CD pipelines and infrastructure automation.
  • Strong troubleshooting and problem-solving skills.
  • Experience working in financial services, fintech, payments, digital banking, or other regulated environments.
Preferred Experience
  • Experience supporting blockchain, digital asset, stablecoin, payment, or fintech platforms.
  • Experience with observability tools such as Grafana, OpenTelemetry, CloudWatch, ELK, Datadog, Splunk, or similar platforms.
  • Experience with Kubernetes and containerized workloads.
  • Experience with Linux administration.
  • Experience with scripting and automation (Python, Bash, PowerShell, or similar).
  • Experience operating or supporting 24x7 production services.
We offer
  • We build a professional & fun working environment.
  • We focus on your growth, yes the long-term growth.
  • We develop the future-ready digital bank platform.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

GFT TECHNOLOGIES SE • Hong Kong

On-site
HKD 700,000 - 1,200,000
Senior SRE: FinTech Platform Reliability & Observability
Senior SRE: FinTech Platform Reliability & Observability

GFT Technologies APAC & GCC • Hong Kong

On-site
HKD 900,000 - 1,500,000
Professional & fun working environment
Growth opportunities
Future-ready digital bank platform
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kody • Hong Kong

On-site
HKD 900,000 - 1,500,000
Competitive Package
A dynamic and innovative team
Collaborative, inclusive working env.
Senior SRE for Fintech Platform Reliability (Hybrid)
Senior SRE for Fintech Platform Reliability (Hybrid)

GFT TECHNOLOGIES SE • Hong Kong

Hybrid
HKD 700,000 - 1,200,000
Senior Infrastructure Developer
Senior Infrastructure Developer

Btse • Hong Kong

On-site
HKD 900,000 - 1,300,000
Competitive compensation package
Team building programs
Associate, Production Support Engineer
Associate, Production Support Engineer

Galaxy • Hong Kong

On-site
HKD 450,000 - 650,000
Health benefits
Headspace coaching
In-office snacks
Site Reliability Engineer - HFT
Site Reliability Engineer - HFT

Selby Jennings • Hong Kong

On-site
HKD 480,000 - 720,000
Senior Site Reliability Engineer (APAC)
Senior Site Reliability Engineer (APAC)

Reap • Hong Kong

On-site
HKD 900,000 - 1,500,000
Senior DevOps Engineer (Trading Platform)
Senior DevOps Engineer (Trading Platform)

Crypto.com • Hong Kong

On-site
HKD 600,000 - 1,100,000
Crypto.com visa card
Flexible work hours
Annual leave
+2
Core DevOps Engineer
Core DevOps Engineer

Selby Jennings • Hong Kong

On-site
HKD 900,000 - 1,300,000
Relocation support