Lead Site Reliability Engineer

Swift Software

Kuala Lumpur

On-site

MYR 120,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Swift is seeking an experienced Site Reliability Engineer in Kuala Lumpur to manage system administration lifecycles, design architectures, automate deployments, and drive reliability across multi-tenant infrastructure.

You will review architectures, implement monitoring, and collaborate with developers to improve reliability. Proficiency in Python scripting and IaC (Ansible, Terraform, Kubernetes, OpenShift) and experience with ITIL, Agile, and cross-border teams are essential.

Qualifications

  • Bachelor’s or Master’s in engineering, CS, IT or equivalent.
  • Minimum 15 years of experience in SRE or sys admin/software development in an international company.
  • Minimum 2 years of experience leading projects.
  • Familiarity with data ingestion using Elastic Search, Logstash, Kibana and Kafka.
  • Experience with CI/CD tools such as Maven, Jenkins, Nexus, Git and Docker.
  • Proficiency in Linux and scripting/automation (Python, PowerShell, YAML).
  • Strong understanding of distributed systems, microservices, REST/SOAP APIs.
  • Experience with ITIL processes and Agile environments.
  • Experience building data-driven metrics and reporting.
  • Strong collaboration across time zones and cultures.

Responsibilities

  • Work through all phases of the system administration life cycle: capacity planning, architecture design, deployment, monitoring and incident management.
  • Develop automation scripts, infrastructure as code, and tooling to improve reliability and reduce manual work.
  • Review architectures, deployment strategies, observability, and ops docs to ensure reliability.
  • Analyze production issues, identify root causes and implement long-term improvements via automation and monitoring.
  • Collaborate with team members and mentor juniors.
  • Ensure efficient handovers with high-quality documentation and training.
  • Automate deployment and operation of multi-tenant infrastructure for resilience and availability.
  • Develop and maintain monitoring tools and dashboards with self-healing capabilities.
  • Participate in on-call rotations and blameless postmortems; drive continuous learning.
  • Work with developers, product teams, and engineers to troubleshoot and integrate reliability improvements.
  • Provide accurate project estimates and adapt plans across project lifecycles.

Skills

SRE
Linux
Python
Automation
CI/CD
Docker
Kubernetes
Ansible
Terraform
OpenShift
REST APIs

Education

Bachelor’s/Master’s in Engineering/CS/IT

Tools

Maven
Jenkins
Nexus
Git
Docker
Kubernetes
OpenShift
Terraform
Ansible
vCenter

Job description

ABOUT US

We’re the world’s leading provider of secure financial messaging services, headquartered in Belgium. We are the way the world moves value – across borders, through cities and overseas. No other organisation can address the scale, precision, pace and trust that this demands, and we’re proud to support the global economy.

We’re unique too. We were established to find a better way for the global financial community to move value – a reliable, safe and secure approach that the community can trust, completely. We’re always striving to be better and are constantly evolving in an ever-changing landscape, without undermining that trust. Five decades on, our vibrant community reflects the complexity and diversity of the financial ecosystem. We innovate diligently, test exhaustively, then implement fast. In a connected and exciting era, our mission has never been more relevant. Swift now has a presence in 200+ countries and legal territories to serve a community of more than 12,000 banks and financial institutions.

Job Description

What to expect:

  • Work through all phases of the system administration life cycle, including capacity planning, architecture design, compliance, deployment & configuration, monitoring, and incident management.
  • Develop automation scripts, infrastructure as code, and tooling using industry best practices to improve system reliability, reduce manual effort, and enable self‑service.
  • Review system architectures design, deployment strategies, observability setups, and operational documentation to ensure reliability and operational excellence.
  • Analyze production issues, identify root causes, and implement long‑term reliability improvements through automation, monitoring, and architectural enhancements.
  • Work collaboratively with other team members and provide guidance to more junior team members.
  • Organize an efficient handover through high quality documentation and training.
  • Automate the deployment and operation of multi‑tenant infrastructure, handling tasks that ensure system resilience and availability.
  • Develop and maintain monitoring tools, dashboards, and self‑healing mechanisms.
  • Participate in on‑call rotations, conduct blameless postmortems, and drive continuous learning.
  • Work closely with developers, product teams, and engineering stakeholders to troubleshoot issues, improve systems, and integrate reliability improvements.
  • Capable of providing accurate project estimates and strategically adapting plans throughout the project lifecycle.
Qualifications
  • Bachelor’s/master’s degree in engineering, Computer Science, IT, or equivalent experience.
  • Minimum 15 years of experience in Site Reliability Engineering or sys admin/software development within an international company.
  • Minimum 2 years of experience leading projects.
  • Familiarity or experience with data ingestion with big data technologies (Elastic Search, Logstash, Kibana and Kafka).
  • Experience with CICD development & deployment tools such as Maven, Jenkins, Nexus, Git, and Docker.
  • Proficiency in Linux OS.
  • Proficiency in scripting and automation (e.g. Python, PowerShell, YAML) with the ability to develop tools and infrastructure as code (Preferably Ansible, Terraform, Kubernetes, OpenShift).
  • Understanding of distributed systems and microservices architectures, including REST and SOAP APIs.
  • Hands‑on experience with ITIL processes, including Incident, Problem, and Continual Improvement is an advantage.
  • Experience working within an Agile‑driven environment.
  • Practical experience in building metrics for data‑driven reporting.
  • Strong interpersonal skills with a customer‑centric mindset and ability to work effectively across diverse cultures.
  • Proven ability to collaborate with both local and remote teams across different time zones.
  • Familiarity with or experience in managing VM hosts using vCenter is an advantage.
What we offer

We give you the freedom to be yourself. We are creating an environment of unique individuals – like you – with different perspectives on the financial industry and the world.

A diverse and inclusive environment in which everyone’s voice counts and where you can reach your full potential. We are committed to an inclusive and accessible recruitment process. If you require a reasonable accommodation related to accessibility during your application or interview, please contact accessibility‑Sysgroup@swift.com or indicate this in your application.

We are proud that what we do has a critical impact on the global financial community and touches almost every aspect of the financial world. Joining Swift gives you unparalleled exposure to knowledge, expertise and technologies.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Associate IT Operations Specialist
Associate IT Operations Specialist

swift • Kuala Lumpur

On-site
MYR 36,000 - 48,000
Senior Customer Support Engineer
Senior Customer Support Engineer

Swift Software • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Associate IT Operations Specialist
Associate IT Operations Specialist

Swift Software • Kuala Lumpur

On-site
MYR 67,000 - 100,000
Lead DevOps Engineer - DevOps Enablement
Lead DevOps Engineer - DevOps Enablement

Swift Software • Kuala Lumpur

On-site
MYR 180,000 - 340,000
Senior Customer Support Engineer
Senior Customer Support Engineer

Swift • Kuala Lumpur

On-site
MYR 70,000 - 110,000
Lead Business Analyst
Lead Business Analyst

Swift • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Senior Business Consultant
Senior Business Consultant

Swift Software • Kuala Lumpur

On-site
MYR 90,000 - 150,000
Technical Advisory Consultant
Technical Advisory Consultant

Swift Software • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Technical Support Engineer - EMEA shift
Technical Support Engineer - EMEA shift

Swift • Kuala Lumpur

On-site
MYR 56,000 - 89,000
Accessibility accommodations
Associate IT Operations Specialist at Swift
Associate IT Operations Specialist at Swift

Swift • Kuala Lumpur

On-site
MYR 89,000 - 179,000
Accessibility accommodation available
Diverse and inclusive environment