Site Reliability Engineer

Momcozy

Singapore

On-site

SGD 120,000 - 180,000

Full time

45 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Momcozy is expanding its overseas backend team in Singapore to own the stability, observability, and automation of the regional production environment. The role involves independent release management, incident response, and close collaboration with HQ teams in China and the regional tech lead.

You will operate AWS resources and Kubernetes clusters across regions, build monitoring dashboards, and contribute to infrastructure-as-code practices, security, and cost optimization.

Qualifications

  • Bachelor's degree or higher in CS/SE or equivalent
  • 5+ years in operations, SRE, or platform engineering
  • Proficient with AWS (multi-region) including EC2, EKS, RDS, IAM, VPC, CloudWatch
  • Strong Kubernetes skills: cluster operations and troubleshooting
  • Experience with Terraform and CI/CD tools (Jenkins / GitLab CI)
  • Linux and networking fundamentals; scripting in Python / Shell / Go
  • Experience with MySQL, Redis, Kafka / RocketMQ and monitoring
  • Familiar with Prometheus / Grafana / ELK stacks
  • Experience handling production incidents and postmortems
  • Legal to work in Singapore; no visa sponsorship

Responsibilities

  • Operate and optimize regional AWS resources and Kubernetes clusters
  • Build and maintain monitoring dashboards, alert rules, and logs
  • Participate in regional incident response and runbooks
  • Build CI/CD pipelines supporting canary releases and fast rollback
  • Own backup, recovery, and performance tuning for regional databases
  • Manage regional IAM and security measures; enforce least-privilege access

Skills

AWS multi-region ops
Kubernetes operations
SRE / platform engineering
Linux fundamentals
Scripting (Python/Shell/Go)
CI/CD practices
Incident response & postmortems
Security & least-privilege access
English proficiency

Education

Bachelor's degree or above in CS/SE

Tools

Terraform
Jenkins
GitLab CI
Prometheus/Grafana
ELK stack
Nginx / API gateways

Job description

Join our overseas backend team to own the stability, observability, and automation of the Singapore production environment. Our smart hardware and mobile applications run across multiple AWS regions, and this role exists to ensure the region has independent release, monitoring, and emergency-response capability. You will hold operational access to the regional production environment, handle incidents alongside the regional tech lead, and collaborate with the HQ infrastructure team on containerization, unified gateway, and middleware governance initiatives.

Who You’ll Work With
  • Reports to: Regional Tech Lead (based in Singapore)
  • Collaborates with: HQ backend engineering team (China), HQ infrastructure team, regional backend engineers, and product/operations teams
  • Team context: Part of a growing regional engineering team; works across time zones with China HQ
Responsibilities
  • Operate and optimize regional AWS resources and Kubernetes clusters, including capacity planning and cost optimization, ensuring environment consistency and configuration traceability
  • Build and maintain monitoring dashboards, alert rules, and log platforms for applications, middleware (MySQL / Redis / Kafka / RocketMQ), and infrastructure, ensuring alerts are accurate and actionable
  • Participate in regional incident response, executing and automating mitigation actions (rollback, scaling, degradation switches); run degradation and failure drills and maintain incident runbooks
  • Build and maintain CI/CD pipelines supporting canary releases and fast rollback; enforce change review and checklists; advance infrastructure-as-code practices
  • Own backup, recovery, tuning, and capacity assessment for regional databases and middleware; partner with developers on slow-query and performance issue resolution
  • Manage regional IAM accounts and access under company policy; support rollout of WAF, gateway, and other security capabilities; enforce least-privilege access and audit trails for data operations
Requirements
Must-Have
  • Bachelor's degree or above in Computer Science, Software Engineering, or a related field, or equivalent industry experience
  • 5+ years of experience in operations, SRE, or platform engineering
  • Proficient with AWS (EC2, ALB, EKS, RDS, ElastiCache, S3, IAM, VPC, CloudWatch), with multi-region operations experience
  • Strong Kubernetes skills: cluster operations, troubleshooting, and resource governance
  • Experience with Terraform or similar infrastructure-as-code tools, and CI/CD tools such as Jenkins / GitLab CI
  • Solid Linux and networking fundamentals; scripting proficiency in at least one of Python / Shell / Go
  • Experience deploying, monitoring, and troubleshooting MySQL, Redis, and Kafka / RocketMQ
  • Familiar with Prometheus / Grafana / ELK or comparable observability stacks
  • Experience handling production incidents and postmortems; able to execute mitigation calmly and methodically under pressure
  • Legally authorized to work in Singapore; this role does not provide visa sponsorship
Nice-to-Have
  • Able to read Java / Spring Boot code and understand application-side issues
  • Hands-on experience with API gateways (Higress / Shenyu / Kong / Nginx) and canary releases
  • Experience with traffic replay, chaos engineering, or failure drills
  • Experience with multi-region, cross-time-zone collaboration
Language
  • Professional working proficiency in English (CEFR B2 or above)
  • Mandarin Chinese is highly advantageous for collaboration with China HQ, but not required
What We Offer
  • Opportunity to build a regional engineering function from the ground up
  • Cross-regional collaboration with a mature HQ engineering team
  • Exposure to large-scale e-commerce and IoT infrastructure
  • Competitive compensation and benefits package
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Momcozy • Singapore

On-site
SGD 120,000 - 180,000
Competitive compensation
Site Reliability Engineer
Site Reliability Engineer

RemotePeople • Singapore

On-site
SGD 120,000 - 190,000
Cloud Site Reliability Engineer
Cloud Site Reliability Engineer

SINGAPORE POOLS (PRIVATE) LIMITED. • Singapore

On-site
SGD 120,000 - 160,000
Total rewards
Health benefits
Learning opportunities
+1
Senior Sql Database Administrator
Senior Sql Database Administrator

momcozy • Singapore

On-site
SGD 120,000 - 170,000
Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

Singapore Pools • Singapore

On-site
SGD 120,000 - 170,000
Total rewards package
Health & wellness benefits
Continuous learning and upskilling
+1
Senior Backend Software Engineer
Senior Backend Software Engineer

Delivery Hero • Singapore

On-site
SGD 80,000 - 120,000
Free food
Health and dental insurance
Learning and development opportunities
SRE Engineer
SRE Engineer

re-zoo-me • Singapore

Hybrid
SGD 90,000 - 130,000
Cloud Engineer Kubernetes
Cloud Engineer Kubernetes

OPENSOURCE TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineering (SRE) Leader
Site Reliability Engineering (SRE) Leader

Patsnap • Singapore

On-site
SGD 180,000 - 300,000
Site Reliability Engineer
Site Reliability Engineer

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 16,000 - 24,000
Health insurance
Professional development
Visa sponsorship