Infrastructure Architect (Head of department)
We are seeking a highly experienced technology leader to take ownership of the company’s overall technology strategy, engineering organization, product development, infrastructure architecture, security, and system reliability.
This is a senior hands-on leadership role combining CTO-level technology strategy with deep Infrastructure / SRE / Cloud Architecture expertise. The successful candidate will define the technical roadmap, lead engineering and SRE teams, architect large-scale distributed systems, and build secure, highly available, multi-cloud infrastructure.
Key Responsibilities
- Lead the company’s overall technology strategy, architecture roadmap, and technical transformation
- Build, lead, mentor, and manage Engineering, DevOps, and SRE teams
- Partner with founders and management on product strategy and technology direction
- Own major architecture decisions, technical standards, and engineering reviews
- Design organization-wide Infrastructure as Code (IaC) and automation standards using Terraform, Ansible, and related technologies
- Architect scalable distributed systems across Java, Python, Go, PHP, and Node.js
- Design microservices, Kubernetes, container orchestration, and service governance strategies
- Lead capacity planning, cloud cost optimization / FinOps, high availability, and scalability
- Design CDN, edge acceleration, traffic routing, multi-CDN scheduling, and global traffic management
- Own enterprise DNS, domain, SSL/TLS certificate lifecycle, WAF, DDoS protection, and security architecture
- Lead DevOps modernization, CI/CD platform development, containerization, and platform engineering
- Design enterprise network and Zero Trust architecture, including firewalls, VPN, proxies, WireGuard, OpenVPN, SD-WAN, and remote-access solutions
- Establish enterprise observability, monitoring, logging, tracing, SLO/SLI, incident management, and reliability standards
- Architect database, storage, backup, replication, sharding, read/write splitting, and multi-active systems
- Design and implement Business Continuity & Disaster Recovery (BCP/DR) strategies with appropriate RTO/RPO targets
- Lead technical incident response and ensure overall system stability, security, availability, and resilience
- Manage technology budgets, infrastructure vendors, and strategic technology partnerships
- Establish engineering best practices, technical governance, development standards, and operational processes
Requirements
- Bachelor's degree in Computer Science, Software Engineering, Information Technology, or a related field
- 8+ years of experience across software engineering, infrastructure, SRE, DevOps, or systems architecture
- 3+ years of experience leading engineering / technology teams
- Strong experience designing large-scale distributed systems and highly available architectures
- Deep expertise in Kubernetes, cloud infrastructure, microservices, DevOps, CI/CD, and Infrastructure as Code
- Strong hands-on experience with Terraform, Ansible, Docker, Kubernetes, and automation frameworks
- Proven experience with AWS, Alibaba Cloud, Tencent Cloud, Huawei Cloud, GCP, Azure, or other major cloud platforms
- Strong understanding of networking, cybersecurity, Zero Trust, WAF, DDoS protection, DNS, CDN, VPN, and enterprise security
- Strong knowledge of database architecture, distributed storage, replication, backup, and disaster recovery
- Experience establishing SRE, observability, SLO/SLI, incident management, and reliability engineering practices
- Strong understanding of backend and application architecture, preferably across Java, Python, Go, PHP, or Node.js
- Strong leadership, communication, problem-solving, and strategic decision-making skills
- Comfortable working in a hands-on technical leadership environment, not purely management
- Prefer candidates from the Klang Valley area. Once the office is set up, candidates will be required to return to the office when necessary or attend meetings at the office.