Jora Malaysia will close on 16th September 2026. Thank you for being with us, we are cheering you on as you continue your career journey.
ModeFair is an expanding tech venture that excels in GovTech. We provide a holistic suite of services, ranging from strategic counsel and conceptualisation to UI/UX design, custom-made development & integration, and routine IT management.
We aim to retain strong talent with a conducive working environment and competitive benefits. Our culture is fast-paced: everyone uses AI heavily in their daily work, delivery speed and efficiency matter, and we hold people accountable for how they use their time. We promote trust, responsibility, and open communication - and we expect high ownership.
About the job
We are hiring a highly hands-on Senior Infrastructure & Cloud lead for our Kuala Lumpur office - someone who can own the function and act as Head of Infrastructure & Cloud. You report to the CTO as the company's principal infrastructure and cloud SME - not a management-only seat.
You personally design, implement, troubleshoot and operate critical cloud, on-premise and production infrastructure while leading infrastructure people, vendors and day-to-day ops. You are accountable for availability, performance, scalability, security and cost, with a target of at least 99.9% monthly service availability.
You own production operations: major incidents, disruptions, capacity planning, production changes, backup, Business Continuity and Disaster Recovery. Strong senior engineers who can own the full Head scope are welcome even if they have not held a Head title before.
Willing to go the extra mile - build IaC, automation, runbooks and ops tooling with Claude Code / AI CLI instead of defaulting to expensive commercial platforms when an in-house build fits - not chat-only.
Full on-site, Monday-Friday at Sunway Tower, Jalan Ampang, Kuala Lumpur. No hybrid or remote. On-call / emergency escalation applies when production is at risk.
Key responsibilities
- Develop and execute the company's cloud, infrastructure and production-operations strategy with the CTO
- Design, operate and continuously improve infrastructure across AWS, Huawei Cloud and on-premise environments
- Personally troubleshoot and resolve complex production, cloud, network and infrastructure issues when required
- Manage Linux systems (Red Hat / Amazon Linux / EulerOS), virtualization, networking, load balancing, DNS, SSL/TLS, storage and application infrastructure
- Manage Cloudflare services including DNS, CDN, WAF, DDoS protection, Waiting Room, Load Balancing and traffic management
- Ensure critical systems achieve or exceed 99.9% monthly service availability
- Establish and operate monitoring, alerting, observability, capacity management and operational dashboards
- Lead P1 and other major production incidents (troubleshooting, stakeholder coordination, recovery, post-incident reviews)
- Conduct capacity planning, infrastructure optimisation and performance tuning
- Support JVM, application and MySQL performance analysis where infrastructure involvement is required
- Develop and maintain Business Continuity and Disaster Recovery (RTO, RPO, backup, replication, failover, failback) and run regular DR / tabletop exercises
- Establish secure CI/CD, infrastructure and production-deployment processes
- Implement Infrastructure as Code and automation (Terraform, Ansible, Bash or Python) - using Claude Code / AI CLI to ship faster where it helps
- Manage production access, privileged accounts, secrets, certificates and infrastructure security controls
- Ensure production changes are reviewed, approved, auditable and reversible
- Work closely with Cybersecurity and Engineering to protect production systems and data
- Manage infrastructure-related cloud and vendor costs; cut unnecessary spend
- Maintain architecture docs, operational procedures and recovery runbooks
- Manage cloud, data-centre, network and infrastructure vendors
- Give the CTO clear visibility of infrastructure risks, incidents, capacity, costs and open operational problems
Skills & experience required
- Strong hands-on experience designing, operating and troubleshooting production infrastructure
- Strong experience with AWS, Huawei Cloud or another major cloud platform
- Strong understanding of TCP/IP, DNS, routing, VPNs, firewalls, load balancing and network troubleshooting
- Experience operating production systems with at least 99.9% SLA requirements - and willingness to be accountable for that target
- Strong monitoring, observability, incident-response and root-cause analysis
- Experience with Business Continuity, Disaster Recovery, backup and recovery testing (including DR / tabletop exercises)
- Experience with Infrastructure as Code, preferably Terraform
- Proficiency in scripting (Bash, Python or similar)
- Experience with CI/CD and production-deployment processes
- Experience with infrastructure security, IAM, privileged access and secrets management
- Strong capacity planning, performance optimisation and cost management
- Hands-on with Claude Code or similar AI engineering / CLI tools for IaC, automation and ops - not chat-only
- Ability to operate independently and stay technically hands-on
- Leadership: experience leading infrastructure, cloud, DevOps, systems or production-ops people, or owning the function as the accountable senior. Strong seniors who have already been the owner may still fit
- Willing to participate in on-call / emergency escalation and support major incidents outside normal hours when needed
Nice to have
- AWS and Huawei Cloud; Cloudflare Business/Enterprise; VMware; RHEL / Amazon Linux; Nginx / Tomcat; JVM / MySQL tuning
- Terraform and Ansible; GitHub Actions; Prometheus / Grafana / PRTG
- Hybrid cloud and high-availability / high-volume production systems
- Relevant certs (AWS, Huawei Cloud, RHCSA/RHCE, VMware VCP, Terraform, networking) - practical results matter more than cert count