Site Reliability Engineering Lead, SRE & Governance, Group Technology

GREENBEEN TECHNOLOGY SERVICES PRIVATE LIMITED

Singapore

On-site

SGD 300,000 - 520,000

Full time

12 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

GREENBEEN TECHNOLOGY SERVICES PRIVATE LIMITED is seeking an SVP, Site Reliability Engineering (SRE) to lead 24/7 infrastructure operations across Hypervisors, OpenShift, Windows, databases, TWS, and Mainframe environments. The role drives resilience, scalability, automation, and operational excellence in hybrid cloud and on-premises settings, aligned to business risk and regulatory expectations.

The role demands 18+ years in IT infrastructure/SRE with proven leadership of large 24/7 teams,

Qualifications

  • 18+ years in IT infrastructure, SRE, or production operations.
  • Leadership of large-scale 24/7 infrastructure teams in banking/financial services.
  • Strong experience in hybrid cloud, data center, and enterprise platforms.
  • Proven ability to align IT with regulatory and risk requirements.

Responsibilities

  • Lead and manage a distributed 24/7 SRE infrastructure team, including shift-based operations and command center functions.
  • Define and execute the SRE strategy aligned to enterprise technology and business priorities.
  • Establish governance across incident, problem, change, release, and capacity management.
  • Drive SLA/SLO/SLI frameworks to ensure service reliability and performance targets.
  • Oversee end-to-end reliability of infrastructure platforms: Cloud & Container (OpenShift/Kubernetes), Hypervisors, private cloud, Windows/Linux, TWS/Mainframe, Databases.
  • Ensure high availability, resilience, and disaster recovery readiness; own lifecycle planning, patching, upgrades, decommissioning.
  • Champion SRE principles: error budgets, toil reduction, automation-first mindset, and end-to-end observability.
  • Lead initiatives to reduce MTTR, incident volume, and manual effort; scale automation.
  • Ensure 24/7 monitoring, incident response, and robust recovery processes; manage major incidents.
  • Embed ITIL best practices across service management processes.
  • Identify infrastructure risks and drive proactive mitigation; ensure regulatory compliance.
  • Partner with security teams on hardening, vulnerability management, and access controls.
  • Collaborate with application, DevOps, security, architecture, and business teams to improve reliability.
  • Provide leadership in large-scale transformation programs (cloud adoption, infra modernization).
  • Mentor senior leaders and establish clear career progression within SRE teams.

Skills

Leadership
Strategic thinking
Crisis management
Executive communication

Education

Bachelor's degree in Computer Science or related field

Tools

OpenShift
Kubernetes
VMware/Hypervisors
Windows
Linux/Unix
TWS/Mainframe
Databases (MariaDB, Postgres, MSSQL, Redis, DB2)
CI/CD & IaC (DevOps)

Job description

Role Summary

The SVP, Site Reliability Engineering (SRE), will lead and oversee the

24/7 infrastructure operations and reliability engineering function

across critical platforms including

Hypervisors (VPC, EPC, OPC), OpenShift , Windows, Databases, TWS, and Mainframe environments. This role is responsible for driving resilience, scalability, automation, and operational excellence across hybrid cloud and on-premises environments, while ensuring alignment with business, risk, and regulatory expectations.

Key Responsibilities
Leadership & Governance
  • Lead and manage a distributed 24/7 SRE infrastructure team, including shift-based operations and command center functions
  • Define and execute the SRE strategy aligned to enterprise technology and business priorities
  • Establish strong governance across incident, problem, change, release, and capacity management
  • Drive SLA/SLO/SLI frameworks to ensure service reliability and performance targets
Infrastructure & Platform Ownership
  • Oversee end-to-end reliability of infrastructure platforms:
  • Cloud & Container: VPC, OpenShift, Kubernetes
  • Compute & Virtualization: Hypervisors (VMware/others), private cloud platforms
  • Enterprise Platforms: Windows, Unix/Linux, TWS, Mainframe, Databases
  • Ensure high availability, resilience, and disaster recovery readiness across all critical systems
  • Own infrastructure lifecycle including capacity planning, patching, upgrades, and decommissioning
Reliability Engineering & Automation
  • Champion SRE principles including error budgets, toil reduction, and automation-first mindset
  • Drive end-to-end observability strategy (monitoring, logging, tracing)
  • Lead initiatives to reduce MTTR, incident volume, and manual operational effort
  • Scale automation across deployment, patching, incident resolution, and self-healing capabilities
Operational Excellence
  • Ensure 24/7 monitoring, incident response, and recovery processes are robust and continuously improved
  • Lead major incident management and command bridge coordination for critical outages
  • Conduct RCA, trend analysis, and preventive engineering improvements
  • Embed ITIL best practices across service management processes
Risk, Compliance & Security
  • Identify infrastructure risks and drive proactive mitigation strategies
  • Ensure compliance with regulatory, audit, and internal security requirements
  • Partner with security teams on hardening, vulnerability management, and access controls
Stakeholder & Cross-Functional Collaboration
  • Collaborate with application, DevOps, security, architecture, and business teams to improve system reliability
  • Provide leadership in large-scale transformation programs (cloud adoption, infra modernization, SRE maturity)
  • Act as a key interface with senior management and external stakeholders
People & Talent Development
  • Build and develop a high-performing SRE organization across L1/L2/L3 layers
  • Drive fungibility, cross-skilling, and leadership development within the team
  • Mentor senior leaders and establish clear career progression frameworks
Requirements
Experience
  • 18+ years of experience in IT infrastructure, SRE, or production operations
  • Proven leadership in managing large-scale 24/7 infrastructure teams in banking/financial services
  • Strong experience in hybrid cloud, data center, and enterprise platforms
Technical Expertise
  • Deep expertise in:
  • Cloud platforms (private/public cloud architectures)
  • Container platforms (OpenShift/Kubernetes)
  • Hypervisors & virtualization technologies
  • Operating systems (Windows, Linux/Unix)
  • Databases (MariaDB, Postgres, MSSQL, Redis, DB2)
  • Enterprise scheduling & legacy systems (TWS, Mainframe)
  • Strong understanding of DevOps, CI/CD, and infrastructure as code
Leadership & Functional Skills
  • Strong strategic thinking with ability to translate business goals into technology outcomes
  • Excellent incident leadership and crisis management skills
  • Proven track record of driving automation and operational transformation
  • Strong stakeholder management and executive communication skills
Other Skills
  • Expertise in ITIL / Service Management frameworks
  • Strong analytical, problem-solving, and decision-making capabilities
  • Ability to manage high-pressure situations and multiple priorities
Key Success Metrics (Optional for your slide/JD refinement)
  • Infrastructure availability (SLA/SLO adherence)
  • Reduction in MTTR / incident volume
  • Automation coverage & reduction in manual toil
  • Capacity utilization and cost optimization
  • Audit and compliance adherence
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SVP, Site Reliability Engineering Lead, SRE & Governance, Group Technology
SVP, Site Reliability Engineering Lead, SRE & Governance, Group Technology

DBS Bank • Singapore

On-site
SGD 300,000 - 520,000
VP, Site Reliability Engineer
VP, Site Reliability Engineer

Ambition Singapore • Singapore

On-site
SGD 300,000 - 520,000
VP, Site Reliability Engineer
VP, Site Reliability Engineer

AMBITION GROUP SINGAPORE PTE. LTD. • Singapore

On-site
SGD 250,000 - 350,000
Service Delivery Lead
Service Delivery Lead

ESPIRE INFOLABS (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Vice President, Site Reliability Engineering (SRE)
Senior Vice President, Site Reliability Engineering (SRE)

Ambition Singapore • Singapore

On-site
SGD 350,000 - 520,000
SL2564 - SRE & Service Delivery Lead
SL2564 - SRE & Service Delivery Lead

FPT Asia Pacific • Singapore

On-site
SGD 90,000 - 130,000
SL2564 - SRE & Service Delivery Lead
SL2564 - SRE & Service Delivery Lead

FPT Asia Pacific Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
VP, Site Reliability Engineer
VP, Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 120,000 - 160,000
Lead Platform Site Reliability Engineer
Lead Platform Site Reliability Engineer

JPMorgan Chase & Co. • Singapore

On-site
SGD 120,000 - 190,000
Site Reliability Engineer - HM: Mukesh
Site Reliability Engineer - HM: Mukesh

NTT Data Singapore • Singapore

On-site
SGD 80,000 - 120,000