DevOps Lead (Platform & Infrastructure Engineering)

CDG ZIG PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CDG ZIG PTE. LTD. is seeking a senior DevOps/Cloud Operations engineer in Singapore to lead production infrastructure, drive platform improvements, and mentor a growing team. The role emphasizes AWS expertise, Terraform stewardship, EKS operations, and secure CI/CD practices.

The successful candidate will optimize cloud spend, advance platform engineering, and ensure reliability across multi-cluster environments with ArgoCD, Bitbucket Runners, and policy-driven automation.

Qualifications

  • 7+ years in DevOps, Cloud Operations, or SRE with leadership experience.
  • Deep production experience in AWS core services at scale.
  • Extensive Terraform codebase maintenance and upgrades.
  • Hands-on with production Kubernetes (EKS) and ArgoCD in multi-environment setups.
  • Proven track record managing self-hosted CI/CD runners (Bitbucket Runners).
  • Proficient scripting for automation (Python/Go/Bash).
  • Experience with modern autoscalers like Karpenter for cost optimization.
  • Awareness of automated compliance, vulnerability management and related tools.

Responsibilities

  • Maintain, refactor, and mature AWS infra and Terraform code with clean remote state and modularity.
  • Lead infrastructure upgrade cycles to minimize debt and security risk.
  • Analyze and optimize cloud spend focusing on compute, storage, and data transfer.
  • Architect transition toward platform engineering patterns for long-term growth.
  • Maintain production EKS clusters with zero-downtime upgrades and optimized node groups.
  • Own and optimize ArgoCD implementation across multi-cluster environments.
  • Tune container workloads with resource requests/limits, autoscaling, and security boundaries.
  • Manage and scale Bitbucket Runners fleet with high availability and efficiency.
  • Maintain Bitbucket Pipelines; embed automated security gates and compliance checks.
  • Enhance full-stack observability: metrics, logs, traces for distributed microservices.
  • Monitor SLAs/SLOs, drive post-mortems and preventative remediation.
  • Mentor DevOps engineers and codify operational excellence and IaC quality.
  • Collaborate with software leads to translate feedback into platform roadmap features.
  • Any ad hoc duties as assigned

Skills

DevOps leadership
Cloud Operations
Site Reliability Engineering
Terraform
Kubernetes (EKS)
ArgoCD
CI/CD pipelines
Bitbucket Runners
Scripting (Python/Go/Bash)
Karpenter
Compliance & security tooling

Education

Bachelor’s degree in a relevant field

Tools

Terraform
Kubernetes (EKS)
ArgoCD
Bitbucket Runners
Trivy
Checkov
OPA
Karpenter
Python
Go
Bash

Job description

Job Responsibilities
  • Maintain, refactor, and mature our existing AWS infrastructure and Terraform codebases, ensuring clean remote state management and modular scalability.
  • Drive proactive infrastructure upgrade cycles (e.g., AWS service deprecations, EKS platform version updates, and Terraform provider upgrades) to minimize technical debt and security risks.
  • Deeply analyze and optimize existing cloud spend—focusing on compute efficiency, storage lifecycles, and cloud-to-cloud/cross-AZ data transfer fees without compromising resilience.
  • Architect the transition of current systems toward next-generation platform engineering patterns (e.g., policy-as-code, self-service infrastructure portals) to support long-term organizational growth.
  • Maintain and optimize running productionAmazon EKSclusters, managing zero-downtime upgrades, node group efficiencies, and ingress/egress configurations.
  • Own and optimize ourArgoCDimplementation, ensuring strict configuration drift detection, clean synchronization policies, and multi-cluster environment alignment.
  • Continuously tune running containerized workloads by refining pod resource requests/limits, autoscaling policies, and cluster security boundaries.
  • Manage, secure, and scale our self-hostedBitbucket Runnersfleet, ensuring high availability, rapid build execution times, and compute cost-efficiency.
  • Maintain and continually optimize existingBitbucket Pipelinesto accelerate build-and-test feedback loops for developers while integrating automated compliance and security gates.
  • Maintain and enhance full-stack observability frameworks (metrics, logs, traces) to ensure deep visibility into distributed microservices.
  • Monitor system SLAs/SLIs/SLOs, eliminate alerting noise, and act as a senior technical escalation point for critical production incidents, driving comprehensive post-mortems and preventative remediation.
  • Guide, upskill, and mentor DevOps engineers, setting high engineering standards for operational tasks, infrastructure-as-code quality, and documentation.
  • Collaborate closely with software development leads to understand operational pain points, translating their feedback into long-term platform roadmap features.
  • Any ad hoc duties as assigned
Job Requirements
  • Preferably 7+ years of dedicated experience in DevOps, Cloud Operations, or Site Reliability Engineering, with at least 2+ years leading engineering teams or overseeing complex production infrastructure environments.
  • Deep production experience maintaining core AWS environments (VPC, EC2, IAM, S3, RDS, Route53, CloudFront) at scale.
  • Strong proficiency in maintaining, refactoring, and upgrading large-scale Terraform codebases (multi-workspace, remote state management).
  • Hands‑on experience operating production Kubernetes (EKS) and driving continuous delivery using ArgoCD in a live, multi‑environment ecosystem is preferred.
  • Proven track record managing and troubleshooting self-hosted CI/CD runner fleets (specifically Bitbucket Runners) and optimizing delivery pipelines.
  • Competency in writing clean, maintainable scripts/tools using Python, Go, or Bash to automate operational runbooks and platform tasks.
  • Experience utilizing modern Kubernetes autoscalers like Karpenter for dynamic, cost‑optimized EKS node provisioning.
  • Experience integrating automated compliance, dependency scanning, and vulnerability management (e.g., Trivy, Checkov, OPA) into existing runtime environments.
  • Certified Kubernetes Administrator (CKA) or AWS Certified DevOps Engineer – Professional will be a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Lead
DevOps Lead

Merqurio • Singapore

On-site
SGD 90,000 - 150,000
Senior DevOps Engineer
Senior DevOps Engineer

StarHub • Singapore

On-site
SGD 90,000 - 130,000
Senior DevOps Engineer
Senior DevOps Engineer

starhub ltd. • Singapore

On-site
SGD 90,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

United States Digital Space LLC • Singapore

On-site
SGD 80,000 - 120,000
Sr. DevOps Engineer (Infrastructure)
Sr. DevOps Engineer (Infrastructure)

ITConnectHK Limited • Singapore

On-site
SGD 80,000 - 120,000
DevOps & Kubernetes Engineer
DevOps & Kubernetes Engineer

Accenture Southeast Asia • Singapore

On-site
SGD 90,000 - 130,000
Cloud Devops Engineer(AWS, Terraform, Kubernetes, Route 53, VMware, IaC, EC2, VPC, IAM, DevOps)
Cloud Devops Engineer(AWS, Terraform, Kubernetes, Route 53, VMware, IaC, EC2, VPC, IAM, DevOps)

NEPTUNEZ SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

Starhub Ltd • Singapore

On-site
SGD 90,000 - 150,000
DevOps SRE
DevOps SRE

Codigo - The Mobile App Company • Singapore

On-site
SGD 70,000 - 100,000
DevOps Engineer
DevOps Engineer

Merquri • Singapore

On-site
SGD 90,000 - 150,000