Senior Site Reliability Engineer

Momcozy

Singapore

On-site

SGD 120,000 - 180,000

Full time

48 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive compensation

Job summary

Momcozy is seeking a Senior Global SRE / Infrastructure Engineer to own and improve reliable infrastructure across AWS and multiple regions. You will lead on-call readiness, design scalable architectures, and actively participate in global architecture initiatives with cross‑time‑zone teams.

You will troubleshoot complex production issues, automate deployments, and help elevate observability, security, and reliability across the international business.

Qualifications

  • Bachelor's degree or equivalent in CS/IT/Networking or related field.
  • 5+ years in SRE/DevOps, cloud infra, or systems engineering.
  • Strong AWS experience and Linux administration.
  • Hands-on Kubernetes and Docker experience.
  • Solid TCP/IP, DNS, DHCP, VLAN, routing, and network troubleshooting.
  • Experience with infrastructure as code (Terraform).
  • CI/CD tooling experience (GitLab CI, Jenkins, Argo CD).
  • Monitoring/observability with Prometheus, Grafana, CloudWatch, or ELK.
  • Scripting in Python or Bash; incident handling and RCA.

Responsibilities

  • Operate and maintain production infrastructure across AWS and multiple regions.
  • Design, deploy, and operate AWS infra (EC2, VPC, ELB, S3, RDS, CloudWatch, IAM).
  • Operate Kubernetes production clusters and Docker workloads; troubleshoot pod/network/storage issues.
  • Define and improve SLI/SLO/SLA; build monitoring and observability systems.
  • Participate in incident response and on-call rotations; perform RCA.
  • Troubleshoot LAN/WAN, VPN, DNS, routing, firewall; connect global offices.
  • Maintain and improve CI/CD pipelines with GitLab, Jenkins, Argo CD.
  • Implement security best practices; IAM, security groups, encryption; compliance.
  • Act as infrastructure owner for local overseas office; coordinate with ISPs/vendors.
  • Document infra; participate in global architecture reviews.

Skills

SRE
DevOps
Cloud infrastructure
Kubernetes
Linux
Networking
CI/CD
Python/Bash scripting
Incident management
High availability

Education

Bachelor's degree in CS/IT/Networking or related field

Tools

AWS
Terraform
GitLab CI
Jenkins
Argo CD
Prometheus
Grafana
CloudWatch
ELK
Docker

Job description

We are looking for a Senior Global SRE / Infrastructure Engineer to help build, operate, and improve reliable infrastructure across multiple geographic regions. This role serves as a key technical owner for infrastructure supporting our international business, with a strong focus on AWS, Kubernetes, Linux, networking, observability, automation, CI/CD, and production reliability. You will independently troubleshoot complex production issues, support regional infrastructure, participate in global architecture initiatives, and continuously improve the reliability, automation, and operational efficiency of our infrastructure.

Who You'll Work With
  • Reports to: Head of Operations & Security (China HQ)
  • Collaborates with: SRE, DBA, DevOps, network, security, and development teams across regions; acts as a technical bridge between overseas offices and the central infrastructure organization
  • Team context: Senior hands‑on role; participates in global architecture reviews and infrastructure projects; cross-time‑zone collaboration
Responsibilities
  • Operate and maintain production infrastructure across AWS and multiple geographic regions, ensuring availability, reliability, performance, and scalability
  • Design, deploy, and operate AWS infrastructure including EC2, VPC, ELB, S3, RDS, CloudWatch, IAM, and related services; support multi-region AWS infrastructure and cloud migration projects
  • Operate and maintain Kubernetes production clusters; manage Docker/containerized workloads and Kubernetes resources; troubleshoot pod, node, networking, storage, and resource-related issues
  • Define and improve SLI, SLO, and SLA for critical services; build and maintain monitoring, alerting, and observability systems using Prometheus, Grafana, CloudWatch, ELK, or comparable tools
  • Participate in production incident response and on-call rotations; lead incident troubleshooting and root cause analysis; develop preventive measures for recurring incidents
  • Troubleshoot LAN/WAN, VPN, DNS, routing, firewall, and Internet connectivity issues; support connectivity between global offices, cloud environments, and business systems
  • Maintain and improve CI/CD platforms and deployment pipelines using GitLab, Jenkins, Argo CD, or similar technologies; automate infrastructure provisioning, deployment, monitoring, and operational tasks
  • Implement infrastructure security best practices; support IAM, security groups, network segmentation, encryption, and access control; ensure infrastructure complies with company security and global data protection requirements
  • Act as the technical infrastructure owner for the local overseas office; support office network, Wi‑Fi, VPN, and IT infrastructure; coordinate with local ISPs, cloud providers, and IT vendors
  • Communicate technical issues and solutions clearly; maintain infrastructure documentation, SOPs, and architecture diagrams; participate in global architecture reviews
Requirements
Must-Have
  • Bachelor's degree or above in Computer Science, Information Technology, Networking, or a related field, or equivalent industry experience
  • 5+ years of experience in SRE, DevOps, cloud infrastructure, or systems engineering
  • Strong hands‑on experience with AWS cloud infrastructure
  • Strong Linux administration and troubleshooting skills
  • Hands‑on experience with Kubernetes and Docker
  • Solid understanding of TCP/IP, DNS, DHCP, VLAN, routing, VPN, load balancing, and network troubleshooting
  • Experience with infrastructure as code, preferably Terraform
  • Experience with CI/CD tools such as GitLab CI, Jenkins, Argo CD, or similar platforms
  • Experience with monitoring and observability tools such as Prometheus, Grafana, CloudWatch, or ELK
  • Strong scripting or programming skills using Python, Bash, or similar languages
  • Experience handling production incidents and performing root cause analysis
  • Strong understanding of high availability, scalability, and disaster recovery
  • Willingness to participate in an on-call rotation for critical production services
  • Legally authorized to work in Singapore; this role does not provide visa sponsorship
Nice-to-Have
  • AWS certifications such as Solutions Architect, DevOps Engineer, or SysOps Administrator
  • Experience with AWS multi‑region architecture
  • Experience with GitOps and Argo CD
  • Experience with Kafka, Redis, Elasticsearch, or other distributed systems
  • Experience with CDN, WAF, DNS, and global traffic management
  • Experience supporting international e‑commerce, IoT, mobile, or SaaS businesses
  • Experience working in multinational or cross‑border organizations
Language
  • Professional working proficiency in English (CEFR B2 or above)
  • Mandarin Chinese is highly advantageous for collaboration with China HQ, but not required
What We Offer
  • Opportunity to own regional infrastructure for a global business
  • Senior technical scope with influence on global architecture decisions
  • Exposure to multi‑region AWS and Kubernetes environments
  • Competitive compensation and benefits package
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Site Reliability Engineer
Cloud Site Reliability Engineer

SINGAPORE POOLS (PRIVATE) LIMITED. • Singapore

On-site
SGD 120,000 - 160,000
Total rewards
Health benefits
Learning opportunities
+1
Senior Sql Database Administrator
Senior Sql Database Administrator

Momcozy • Singapore

On-site
SGD 120,000 - 180,000
Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

Singapore Pools • Singapore

On-site
SGD 120,000 - 170,000
Total rewards package
Health & wellness benefits
Continuous learning and upskilling
+1
Senior Site Reliability Engineer / SRE Lead
Senior Site Reliability Engineer / SRE Lead

Reolink Technology Pte. Ltd. • Singapore

On-site
SGD 120,000 - 180,000
Insurance Coverage
Yearly Bonus & Performance Bonus
SVP, Site Reliability Engineering Lead, SRE & Governance, Group Technology
SVP, Site Reliability Engineering Lead, SRE & Governance, Group Technology

DBS Bank • Singapore

On-site
SGD 300,000 - 520,000
Site Reliability Engineer( SRE)
Site Reliability Engineer( SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
DevOps Engineer – Investment Technology (SG/HK)
DevOps Engineer – Investment Technology (SG/HK)

Io Tech Solutions Limited • Singapore

On-site
SGD 90,000 - 140,000
Site Reliability Engineer (SRE) - Multiple Opportunities, Tech-driven and Innovative Environment
Site Reliability Engineer (SRE) - Multiple Opportunities, Tech-driven and Innovative Environment

DADACONSULTANTS PTE. LTD. • Singapore

On-site
SGD 60,000 - 120,000
Site Reliability Engineering (SRE) Leader
Site Reliability Engineering (SRE) Leader

Patsnap • Singapore

On-site
SGD 180,000 - 300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

landi international (singapore) pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000