Cloud Operations Engineer – Infrastructure

TP-LINK CORPORATION PTE. LTD.

Singapore

On-site

SGD 110,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TP-LINK CORPORATION PTE. LTD. in Singapore seeks a hands-on Platform Engineer to design, build and operate reliable cloud-native infrastructure for large-scale workloads.

You will manage multi-account AWS environments, production Kubernetes clusters and GitOps workflows to ensure security, scalability and auditable deployments. The role emphasizes collaboration with application engineering, security and architecture teams, on-call rotation, and continuous improvement of reliability practices

Qualifications

  • Bachelor’s degree in Computer Science, Software Engineering, Information Technology or related discipline.
  • At least two years of hands-on experience in cloud infrastructure, Kubernetes operations, platform engineering, SRE or DevOps.
  • Strong knowledge of AWS, particularly EKS, IAM, VPC, EC2, S3 and security capabilities.
  • Hands‑on experience operating Kubernetes in production: cluster provisioning, workloads, networking, autoscaling, monitoring.

Responsibilities

  • Design, build and maintain reliable, scalable and secure cloud-native infrastructure for large-scale production workloads.
  • Operate and optimise multi-account AWS environments, ensuring infrastructure is secure, repeatable and auditable.
  • Provision, administer and upgrade production Kubernetes clusters.
  • Manage Kubernetes networking, autoscaling, observability, capacity planning and day-to-day cluster operations.
  • Build and operate Kubernetes ecosystem components (CRDs, Helm, HPA, Cluster Autoscaler, CoreDNS, Cluster API).
  • Develop and improve GitOps deployment workflows using FluxCD, ArgoCD or similar tools.
  • Manage and enhance Istio service mesh capabilities (traffic routing, service discovery, resilience, security).
  • Automate infrastructure provisioning and configuration management using Terraform and IaC tools.
  • Improve CI/CD pipelines and observability platforms using Go, Python or similar.
  • Establish reliability practices: SLOs, incident response, monitoring and post-incident reviews.
  • Investigate and resolve complex production issues across AWS, Kubernetes, Linux and networking.
  • Collaborate with engineering, security and architecture teams to improve reliability and scalability.
  • Participate in scheduled on-call rotation to support production cloud infra and Kubernetes.

Skills

Cloud infrastructure
Kubernetes operations
Site Reliability Eng
DevOps practices
Go
Python
GitOps
Monitoring
Incident response
On-call readiness

Education

Bachelor’s degree

Tools

Kubernetes
Helm
FluxCD
ArgoCD
Terraform
Cluster API
CoreDNS
HPA
Cluster Autoscaler
AWS
CI/CD tooling

Job description

Key Responsibilities
  • Design, build and maintain reliable, scalable and secure cloud-native infrastructure supporting large-scale production workloads.

  • Operate and optimise multi-account AWS environments, ensuring infrastructure is secure, repeatable and auditable.

  • Provision, administer and upgrade production Kubernetes clusters.

  • Manage Kubernetes networking, autoscaling, observability, capacity planning and day-to-day cluster operations.

  • Build and operate Kubernetes ecosystem components, including CRDs, Helm, HPA, Cluster Autoscaler, CoreDNS and Cluster API.

  • Develop and improve GitOps deployment workflows using FluxCD, ArgoCD or similar tools.

  • Manage and enhance Istio service mesh capabilities, including traffic routing, service discovery, resilience, security and service-to-service communication.

  • Automate infrastructure provisioning and configuration management using Terraform and related Infrastructure as Code tools.

  • Improve CI/CD pipelines, observability platforms and operational workflows using Go, Python or other appropriate technologies.

  • Establish and continuously improve reliability practices, including Service Level Objectives, Error Budgets, monitoring, alerting, incident response and post-incident reviews.

  • Investigate and resolve complex production issues involving AWS infrastructure, Kubernetes, Linux, networking and distributed services.

  • Collaborate with application engineering, architecture, security and platform teams to improve system reliability, scalability and operational efficiency.

  • Participate in a scheduled on-call rotation supporting production cloud infrastructure and Kubernetes platforms.

Requirements
  • Bachelor’s degree in Computer Science, Software Engineering, Information Technology or a related discipline.

  • At least two years of hands-on experience in cloud infrastructure, Kubernetes operations, platform engineering, Site Reliability Engineering, DevOps or a related area.

  • Strong knowledge of AWS, particularly EKS, IAM, VPC, EC2, S3 and associated networking and security capabilities.

  • Hands‑on experience operating Kubernetes in a production environment, including:

    • Cluster provisioning and architecture

    • Workload orchestration

    • Networking and service discovery

    • Autoscaling and capacity management

    • Monitoring and troubleshooting

  • Familiarity with Kubernetes ecosystem tools such as CRDs, Helm, Cluster API, HPA, Cluster Autoscaler and CoreDNS.

  • Experience with GitOps tools such as FluxCD or ArgoCD.

  • Solid Linux administration and troubleshooting skills, including systemd, networking and performance analysis.

  • Experience with CI/CD pipelines and infrastructure automation using Terraform, Go, Python or similar tools.

  • Good understanding of reliability engineering practices, including SLOs, incident response, monitoring, alerting and post‑incident reviews.

  • Strong analytical and problem‑solving skills, with the ability to diagnose complex infrastructure issues across distributed systems.

  • Good communication and collaboration skills, with the ability to work effectively across engineering, security and architecture teams.

  • Willingness to participate in a scheduled on‑call rotation.

Additional Advantages
  • Experience with NVIDIA device plugins, GPU scheduling or GPU workload operations in Kubernetes.

  • Experience with other public cloud platforms, particularly Azure or Alibaba Cloud.

  • Kubernetes certifications such as CKA, CKAD or CKS.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Software Engineer
Cloud Software Engineer

United States Digital Space LLC • Singapore

On-site
SGD 70,000 - 100,000
Cloud Manager
Cloud Manager

KWE-APLL Technology Services Pte Ltd • Singapore

On-site
SGD 180,000 - 240,000
Cloud Engineering
Cloud Engineering

ITCAN Pte Limited • Singapore

On-site
SGD 120,000 - 180,000
Platform Operations Engineer (Cloud Infrastructure)
Platform Operations Engineer (Cloud Infrastructure)

TEKISHUB CONSULTING SERVICES PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Cloud Engineer
Cloud Engineer

Exasoft Pte. Ltd. • Singapore

On-site
SGD 60,000 - 90,000
AWS Clouds Operation Engineer
AWS Clouds Operation Engineer

QUESS SELECTION & SERVICES PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Cloud Devops Engineer(AWS, Terraform, Kubernetes, Route 53, VMware, IaC, EC2, VPC, IAM, DevOps)
Cloud Devops Engineer(AWS, Terraform, Kubernetes, Route 53, VMware, IaC, EC2, VPC, IAM, DevOps)

NEPTUNEZ SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
DevOps Lead (Platform & Infrastructure Engineering)
DevOps Lead (Platform & Infrastructure Engineering)

CDG ZIG PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Infrastructure Engineer – IDC, Bare-Metal Kubernetes
Senior Infrastructure Engineer – IDC, Bare-Metal Kubernetes

Jobtailor • Singapore

On-site
SGD 120,000 - 190,000
Senior DevOps Engineer
Senior DevOps Engineer

starhub ltd. • Singapore

On-site
SGD 90,000 - 180,000