Sr. SRE

United States Digital Space LLC

Singapore

On-site

SGD 120,000 - 180,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

On-site in Singapore (3 days/wk)

Job summary

United States Digital Space LLC is seeking a Sr. Site Reliability Engineer to reinforce release management, CI/CD pipelines, and infrastructure stability across development, staging, and production environments.

You will work with cloud and on‑premises infrastructure (AWS, GCP) using Terraform, Ansible, CloudFormation, Docker, Kubernetes, and observability tools to ensure high availability, security, and compliant operations while driving automation and reliability improvements.

Qualifications

  • Bachelor's degree in a relevant field; 8+ years of platform‑infrastructure experience.
  • Hands‑on ownership of CI‑CD and cloud operations; Kubernetes expertise.
  • Proficient in scripting (Python, Bash, PowerShell) and IaC tooling.

Responsibilities

  • Support end‑to‑end release management and deployment coordination.
  • Build, maintain, and optimize CI‑CD pipelines for automated builds and deployments.
  • Troubleshoot pipeline failures and deployment blockers to ensure reliable releases.
  • Collaborate with development teams to improve release efficiency and reduce risk.
  • Manage cloud infrastructure (AWS, GCP) and on‑prem environments for high availability.

Skills

Platform‑infrastructure
CI-CD
Kubernetes
Terraform
Scripting
Incident management

Education

Bachelor's degree in a relevant field
Master's degree or higher (preferred)

Tools

Terraform
Ansible
CloudFormation

Job description

About Us

the company is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid.


At the company, you'll have the opportunity to create impact at scale — tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world.


Join the company and do work that matters – to you, to your community, and to the world. Progress starts with you.


Job Description

The Sr. Site Reliability Engineer is responsible for supporting release management, CI-CD pipeline operations, infrastructure stability, and server lifecycle maintenance across development, staging, and production environments. This role focuses on ensuring reliable, secure, and efficient delivery of applications through robust automation, infrastructure management, and operational best practices.


The engineer works closely with development and platform teams to build, maintain, and optimize CI-CD pipelines, enabling automated build, test, and deployment processes. They contribute to release planning, coordination, and execution, ensuring successful and predictable production releases.


This role provides hands‑on support for cloud and on‑premise infrastructure (AWS, GCP), including server provisioning, patching, upgrades, and ongoing maintenance to ensure high availability, security, and compliance. The engineer applies infrastructure-as-code practices (Terraform, Ansible, CloudFormation) to standardize and scale environment management.


Additionally, the engineer supports containerized services (Docker, Kubernetes), assists in troubleshooting production issues, and ensures operational readiness through system monitoring and performance optimization. While observability tooling is utilized, the focus remains on operational reliability, release execution, and infrastructure health.


All roles require digital fluency, including the ability to leverage emerging technologies such as Generative AI tools (e.g., Claude, ChatGPT, Microsoft Copilot) to improve productivity and streamline workflows.


Key Responsibilities:


Release Management & CI-CD:



  • Support end-to-end release management, including planning, coordination, and validation of production deployments.

  • Build, maintain, and optimize CI-CD pipelines for automated build, test, and deployment.

  • Troubleshoot and resolve pipeline failures, deployment issues, and release blockers.

  • Partner with development teams to improve release efficiency and reduce deployment risk.


Infrastructure & Server Lifecycle Management:



  • Manage and support cloud infrastructure (AWS, GCP) ensuring availability, reliability, and security.

  • Perform server patching, upgrades, and maintenance in line with security and compliance requirements.

  • Provision and manage infrastructure using infrastructure-as-code tools (Terraform, Ansible, CloudFormation).

  • Ensure environment consistency across development, staging, and production.


Containerized Service Support:



  • Support implementation and operations of containerized services (Docker, Kubernetes).

  • Maintain platform stability and optimize performance of hosted applications.


Operations & Incident Support:



  • Monitor system health and performance, proactively identifying and resolving issues.

  • Participate in incident response, troubleshooting, and root cause analysis for production systems.

  • Provide first-level support for infrastructure and deployment issues, escalating as needed.


Scripting & Automation:



  • Develop and maintain automation scripts to streamline infrastructure operations, deployments, and routine maintenance tasks.

  • Use scripting languages such as Python, Bash, or PowerShell to improve operational efficiency and reduce manual intervention.

  • Automate server patching, environment provisioning, health checks, and deployment workflows.

  • Integrate scripts into CI-CD pipelines to support end-to-end automation.

  • Collaborate with teams to identify automation opportunities and standardize reusable scripts and tooling.


Automation & Continuous Improvement:



  • Identify opportunities to automate repetitive operational tasks and improve workflows.

  • Drive improvements in deployment processes, infrastructure reliability, and operational efficiency.


Documentation & Standards:



  • Create and maintain documentation for release processes, infrastructure configurations, and operational procedures.

  • Ensure adherence to operational and security best practices.


the company requires at least 3 days in office, expectations of these days will be confirmed by your Hiring Manager.


Qualifications

Education & Experience:



  • Bachelor's degree in a relevant field plus 8+ years of relevant work experience, OR

  • Advanced degree (Master's, MBA, etc.) plus 5+ years of relevant work experience, OR

  • PhD plus 2+ years of relevant work experience, OR

  • 11+ years of relevant work experience without the degree path specified

  • A strong candidate would typically have 8+ years of platform‑infrastructure experience, proven ownership of CI‑CD and cloud operations, and hands‑on expertise with Kubernetes, Terraform, scripting, and production incident management


the company is an EEO Employer

Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability or protected veteran status. the company will also consider for employment qualified applicants with criminal histories in a manner consistent with EEOC guidelines and applicable local law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. SRE
Sr. SRE

VISA WORLDWIDE PTE. LIMITED • Singapore

On-site
SGD 120,000 - 180,000
Sr. SRE
Sr. SRE

Visa • Singapore

On-site
SGD 140,000 - 210,000
Software Engineer/ Site Reliability Engineer
Software Engineer/ Site Reliability Engineer

United States Digital Space LLC • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 200,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

VANGUARD SOFTWARE PTE. LTD. • Singapore

On-site
SGD 100,000 - 150,000
Technical Leadership
Career Growth
High-Performance Team
+1
Site Reliability Engineer(Senior SRE)
Site Reliability Engineer(Senior SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

United States Digital Space LLC • Singapore

On-site
SGD 100,000 - 130,000
Site Reliability Engineer (Splunk, Python, OCI, Dynatrace, RCA, Terraform, Ansible )
Site Reliability Engineer (Splunk, Python, OCI, Dynatrace, RCA, Terraform, Ansible )

NEPTUNEZ SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Software Engineer/ Site Reliability Engineer
Software Engineer/ Site Reliability Engineer

Visa • Singapore

On-site
SGD 120,000 - 180,000
Senior Site Reliability Engineer - Cloud & CI/CD
Senior Site Reliability Engineer - Cloud & CI/CD

Visa • Singapore

On-site
SGD 140,000 - 210,000