DevOps Lead (Platform & Infrastructure Engineering)

CDG ZIG PTE. LTD.

Singapore

On-site

SGD 140,000 - 190,000

Full time

12 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

CDG ZIG PTE. LTD. is seeking a Senior DevOps professional to mature our AWS/EKS infrastructure and scale our platform engineering practices.

You will own Terraform codebases, drive upgrades and cost optimizations while maintaining high-availability production clusters and ArgoCD pipelines. Lead and mentor a team of DevOps engineers, collaborating with software leads to implement policy-as-code, self-service portals, and secure, scalable workflows.

Qualifications

  • 7+ years in DevOps, Cloud Operations or SRE with 2+ years leading teams
  • Deep AWS experience (VPC, EC2, IAM, S3, RDS, Route53, CloudFront) at scale
  • Strong Terraform knowledge (multi-workspace, remote state)
  • Hands-on with production Kubernetes (EKS) and ArgoCD in multi-environment setups
  • Experience managing self-hosted CI/CD runner fleets (Bitbucket Runners) and pipelines
  • Scripting in Python/Go/Bash to automate runbooks and tasks
  • Experience with modern autoscalers like Karpenter for cost-optimized EKS
  • Integration of automated compliance, dependency scanning, vulnerability mgmt (Trivy, Checkov, OPA)
  • CKA or AWS DevOps Pro certification a plus

Responsibilities

  • Maintain and mature AWS infrastructure and Terraform codebases with clean remote state
  • Drive infrastructure upgrades to minimize debt and risks
  • Analyze and optimize cloud spend for compute, storage, and inter-region data transfer
  • Architect transition to next-gen platform engineering patterns
  • Maintain production EKS clusters with zero-downtime upgrades
  • Own and optimize ArgoCD with drift detection and multi-cluster alignment
  • Tune container workloads with resource requests/limits, autoscaling, and security boundaries
  • Manage Bitbucket Runners fleet for high availability and cost efficiency
  • Maintain Bitbucket Pipelines and integrate automated compliance gates
  • Maintain observability with metrics, logs, traces for distributed microservices
  • Monitor SLAs/SLIs/SLOs; drive post-mortems and remediation
  • Mentor DevOps engineers and elevate platform engineering standards
  • Collaborate with software leads to translate feedback into roadmaps
  • Leverage Generative AI to redesign workflows and reduce repetitive work
  • Handle ad hoc duties as assigned

Skills

AWS
Terraform
Kubernetes (EKS)
ArgoCD
Bitbucket Runners
Python/Go/Bash
CI/CD pipelines
Cost optimization
Automation scripting
Security/compliance tooling

Tools

Karpenter
Trivy
Checkov
OPA

Job description

Job Responsibilities:
  • Maintain, refactor, and mature our existing AWS infrastructure and Terraform codebases, ensuring clean remote state management and modular scalability.
  • Drive proactive infrastructure upgrade cycles (e.g., AWS service deprecations, EKS platform version updates, and Terraform provider upgrades) to minimize technical debt and security risks.
  • Deeply analyze and optimize existing cloud spend—focusing on compute efficiency, storage lifecycles, and cloud‑to‑cloud/cross‑AZ data transfer fees without compromising resilience.
  • Architect the transition of current systems toward next‑generation platform engineering patterns (e.g., policy‑as‑code, self‑service infrastructure portals) to support long‑term organizational growth.
  • Maintain and optimize running production Amazon EKS clusters, managing zero‑downtime upgrades, node group efficiencies, and ingress/egress configurations.
  • Own and optimize our ArgoCD implementation, ensuring strict configuration drift detection, clean synchronization policies, and multi‑cluster environment alignment.
  • Continuously tune running containerized workloads by refining pod resource requests/limits, autoscaling policies, and cluster security boundaries.
  • Manage, secure, and scale our self‑hosted Bitbucket Runners fleet, ensuring high availability, rapid build execution times, and compute cost‑efficiency.
  • Maintain and continually optimize existing Bitbucket Pipelines to accelerate build‑and‑test feedback loops for developers while integrating automated compliance and security gates.
  • Maintain and enhance full‑stack observability frameworks (metrics, logs, traces) to ensure deep visibility into distributed microservices.
  • Monitor system SLAs/SLIs/SLOs, eliminate alerting noise, and act as a senior technical escalation point for critical production incidents, driving comprehensive post‑mortems and preventative remediation.
  • Guide, upskill, and mentor DevOps engineers, setting high engineering standards for operational tasks, infrastructure‑as‑code quality, and documentation.
  • Collaborate closely with software development leads to understand operational pain points, translating their feedback into long‑term platform roadmap features.
  • Use Generative AI and automation to redesign workflows, reduce repetitive work and human error, improve decision quality, and build the team's strategic consulting capability
  • Any ad hoc duties as assigned
Job Requirements:
  • Preferably 7+ years of dedicated experience in DevOps, Cloud Operations, or Site Reliability Engineering, with at least 2+ years leading engineering teams or overseeing complex production infrastructure environments.
  • Deep production experience maintaining core AWS environments (VPC, EC2, IAM, S3, RDS, Route53, CloudFront) at scale.
  • Strong proficiency in maintaining, refactoring, and upgrading large‑scale Terraform codebases (multi-workspace, remote state management).
  • Hands‑on experience operating production Kubernetes (EKS) and driving continuous delivery using ArgoCD in a live, multi‑environment ecosystem is preferred.
  • Proven track record managing and troubleshooting self-hosted CI/CD runner fleets (specifically Bitbucket Runners) and optimizing delivery pipelines.
  • Competency in writing clean, maintainable scripts/tools using Python, Go, or Bash to automate operational runbooks and platform tasks.
  • Experience utilizing modern Kubernetes autoscalers like Karpenter for dynamic, cost‑optimized EKS node provisioning.
  • Experience integrating automated compliance, dependency scanning, and vulnerability management (e.g., Trivy, Checkov, OPA) into existing runtime environments.
  • Certified Kubernetes Administrator (CKA) or AWS Certified DevOps Engineer – Professional will be a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Lead (Platform & Infrastructure Engineering)
DevOps Lead (Platform & Infrastructure Engineering)

Zig by ComfortDelGro • Singapore

On-site
SGD 90,000 - 130,000
DevOps Engineer
DevOps Engineer

IO TECH SOLUTIONS LIMITED • Singapore

On-site
SGD 90,000 - 150,000
Devops
Devops

Cognizant • Singapore

On-site
SGD 120,000 - 180,000
Senior DevOps Engineer
Senior DevOps Engineer

starhub ltd. • Singapore

On-site
SGD 90,000 - 180,000
DevOps Engineer
DevOps Engineer

BROADCHAINS FINTECH PTE. LTD. • Singapore

On-site
SGD 100,000 - 140,000
Lead DevSecOps Engineer: Cloud, Kubernetes & CI/CD
Lead DevSecOps Engineer: Cloud, Kubernetes & CI/CD

KrisShop Pte. Ltd. • Singapore

On-site
SGD 90,000 - 150,000
DevOps Engineer
DevOps Engineer

Argyll Scott • Singapore

On-site
SGD 120,000 - 180,000
DevOps Engineer - #1589
DevOps Engineer - #1589

JOBSTER PRIVATE LTD. • Singapore

On-site
SGD 90,000 - 150,000
Lead, DevSecOps
Lead, DevSecOps

KrisShop Pte. Ltd. • Singapore

On-site
SGD 90,000 - 150,000
Platform Infrastructure Engineer
Platform Infrastructure Engineer

Cognizant • Singapore

On-site
SGD 120,000 - 180,000
Remote work option