Established global technology company with a strong international presence and large-scale commercial operations
Growing its technology team in Singapore, supporting business-critical platforms across multiple markets
Operating large-scale cloud infrastructure, with strong focus on reliability, security and operational excellence
We are hiring for both Cloud SRE and Infrastructure SRE in Singapore
Key Responsibilities
- Manage and maintain production cloud infrastructure, primarily on AWS
- Ensure the reliability, availability and performance of critical production systems
- Manage infrastructure provisioning and configuration through Infrastructure-as-Code (IaC)
- Support containerised environments and Kubernetes-based infrastructure
- Monitor production systems, troubleshoot incidents and conduct root cause analysis
- Manage Linux-based production environments and perform system-level troubleshooting
- Support database operations including backup, recovery, performance and access management
- Maintain network configurations, firewall policies and infrastructure access controls
- Implement security patches and address infrastructure vulnerabilities within required timelines
- Support deployment, release and infrastructure change management processes
- Build and maintain monitoring, alerting and observability capabilities
- Maintain technical documentation, infrastructure diagrams and operational runbooks
- Work closely with software engineering and other technical teams on production changes and infrastructure requirements
Requirements
- Bachelor's degree in Computer Science, Information Technology or a related discipline
- 3+ years of experience in SRE, DevOps, Cloud Infrastructure, Systems Engineering or a similar role
- Solid knowledge of Linux systems administration and troubleshooting
- Hands-on experience with Terraform, CloudFormation, Ansible or similar IaC/configuration tools
- Experience with Kubernetes and containerised environments
- Familiarity with CI/CD tools such as Jenkins, GitLab CI or GitHub Actions
- Experience with monitoring and observability platforms such as Prometheus, Grafana, ELK or CloudWatch
- Strong troubleshooting skills and experience handling production incidents
- Experience in AWS will be a good plus
Regrettably, only shortlisted candidates will be notified.