Cloud Engineer Kubernetes

OPENSOURCE TECHNOLOGIES PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

33 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

OPENSOURCE TECHNOLOGIES PTE. LTD. in Singapore is seeking experienced Kubernetes and Site Reliability Engineers to support highly scalable production platforms for a global technology customer.

You will apply hands-on expertise in Kubernetes, Linux, automation, observability, incident management and troubleshooting of distributed systems. The role requires comfort operating large-scale production environments with high availability, performance, and operational excellence.

Qualifications

  • Hands-on Kubernetes administration and troubleshooting.
  • Strong incident management and RCA experience.
  • Experience automating routine operations with Python, Bash or similar scripting.
  • Understanding of SRE principles and reliability metrics.
  • Familiarity with monitoring/observability tools (Prometheus, Grafana, ELK/OpenSearch).
  • Experience with container platforms and cloud/private-cloud infrastructure.

Responsibilities

  • Operate, maintain and troubleshoot large-scale Kubernetes production environments.
  • Ensure reliability, scalability, availability and performance of critical services.
  • Investigate complex production issues and perform root-cause analysis.
  • Participate in incident response and drive permanent corrective actions.
  • Automate repetitive operational activities and improve platform reliability.
  • Build and improve monitoring, alerting, logging and observability frameworks.
  • Define and track SLIs, SLOs and operational reliability metrics.
  • Support Kubernetes upgrades, patching and platform improvements.
  • Collaborate with engineering, infrastructure, security and DevOps teams.
  • Perform capacity planning, performance tuning and reliability improvements.
  • Develop runbooks and automation scripts for operations.
  • Participate in production readiness reviews and ensure operational standards.

Skills

Incident management
Root-cause analysis
Automation scripting
Observability principles
Performance tuning
Distributed systems
Capacity planning
Team collaboration

Tools

Kubernetes
Linux
Prometheus
Grafana
Splunk/OpenSearch/ELK
Datadog
Helm
Argo CD
Flux
GitOps
Python
Bash

Job description

Role Overview
  • We are looking for experienced Kubernetes & Site Reliability Engineers to support highly scalable, business-critical production platforms for a global technology customer in Singapore.
  • The role requires strong hands-on expertise in Kubernetes, Linux, production reliability, automation, observability, incident management and troubleshooting of distributed systems
  • Candidates should be comfortable operating large-scale production environments where availability, performance, automation and operational excellence are critical.
Key Responsibilities
  • Operate, maintain and troubleshoot large-scale Kubernetes-based production environments
  • Ensure reliability, scalability, availability and performance of critical services.
  • Investigate complex production issues and perform detailed root-cause analysis.
  • Participate in incident response and drive permanent corrective actions.
  • Automate repetitive operational activities and improve platform reliability.
  • Build and improve monitoring, alerting, logging and observability frameworks.
  • Define and track SLIs, SLOs and operational reliability metrics
  • Support Kubernetes upgrades, configuration changes, patching and platform improvements.
  • Work closely with application engineering, infrastructure, platform, security and DevOps teams.
  • Perform capacity planning, performance tuning and reliability improvements.
  • Develop and maintain operational runbooks, automation scripts and troubleshooting documentation.
  • Participate in production readiness reviews and ensure applications meet operational standards.
Mandatory Skills
  • Strong hands-on experience with
  • Kubernetes administration and troubleshooting
  • Strong understanding of Kubernetes architecture, including:
  • Pods
  • Deployments
  • StatefulSets
  • Services
  • Ingress
  • ConfigMaps / Secrets
  • RBAC
  • Storage
  • Networking
  • Strong
  • Linux systems administration and troubleshooting skills.
  • Good understanding of networking concepts such as DNS, TCP/IP, load balancing and service connectivity.
  • Strong understanding of
  • Site Reliability Engineering principles
  • Experience supporting large-scale, high-availability production systems.
  • Strong incident management and RCA experience.
  • Hands-on scripting/automation experience using
  • Python, Bash/Shell or similar
  • Experience with monitoring and observability tools such as
  • Prometheus, Grafana, Splunk, ELK/OpenSearch, Datadog or equivalent
  • Helm or similar Kubernetes package/deployment management tools.
  • GitOps experience using tools such as Argo CD or Flux.
  • Knowledge of service mesh concepts.
  • Experience with container security and Kubernetes security practices.
  • Experience with cloud or private-cloud infrastructure.
  • Familiarity with distributed systems and microservices architectures.
  • Exposure to performance engineering and capacity management.
  • Experience working in globally distributed engineering environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer / Site Reliability Engineer (Kubernetes)
Platform Engineer / Site Reliability Engineer (Kubernetes)

BOUNTEOUSXACCOLITE SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 170,000
Kubernetes Platform Engineer – Global Investment Firm
Kubernetes Platform Engineer – Global Investment Firm

Pinpoint Asia • Singapore

On-site
SGD 120,000 - 180,000
Cloud Automation & Kubernetes Specialist
Cloud Automation & Kubernetes Specialist

EAMES CONSULTING GROUP (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Platform Engineer / Specialist (Kubernetes)
Platform Engineer / Specialist (Kubernetes)

THIRD PARTY CONSULTING PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Kubernetes Support Engineer / Site Resident Engineer (Cloud Native)
Kubernetes Support Engineer / Site Resident Engineer (Cloud Native)

re-zoo-me • Singapore

Hybrid
SGD 60,000 - 100,000
Cloud Operations Engineer – Infrastructure
Cloud Operations Engineer – Infrastructure

TP-LINK CORPORATION PTE. LTD. • Singapore

On-site
SGD 110,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kidentify • Singapore

On-site
SGD 120,000 - 180,000
Lead Platform Engineer (Azure, Kubernetes & DevOps)
Lead Platform Engineer (Azure, Kubernetes & DevOps)

Hudson Singapore • Singapore

On-site
SGD 120,000 - 180,000
null
Cloud Operations Engineer Engineering and Technology Singapore Experienced (Individual Contributor) Infrastructure
Cloud Operations Engineer Engineering and Technology Singapore Experienced (Individual Contributor) Infrastructure

SEA Singapore • Singapore

On-site
SGD 60,000 - 90,000
Support Engineer - Kubernetes (K8s)
Support Engineer - Kubernetes (K8s)

MICHAEL PAGE (PERSONNEL) PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000