Kubernetes & Site Reliability Engineer (SRE)

OPENSOURCE TECHNOLOGIES PTE. LTD.

Singapore

On-site

SGD 90,000 - 130,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

OPENSOURCE TECHNOLOGIES PTE. LTD. in Singapore seeks experienced Kubernetes and Site Reliability Engineers to support highly scalable production platforms for a global technology customer.

You will work on Kubernetes administration, incident management, automation, and observability to ensure reliability and performance. Responsibilities include operating large-scale clusters, capacity planning, and collaborating with cross-functional teams.

Qualifications

  • Strong hands-on experience with Kubernetes administration and troubleshooting.
  • Solid Linux systems administration and troubleshooting skills.
  • Good understanding of networking concepts (DNS, TCP/IP, load balancing, service connectivity).
  • Site Reliability Engineering principles and incident management experience.
  • Hands-on scripting/automation using Python or Bash/Shell.
  • Experience with monitoring/observability tools: Prometheus, Grafana, Splunk, ELK/OpenSearch, Datadog.
  • Helm or similar Kubernetes packaging tools; GitOps with Argo CD/Flux.
  • Knowledge of service mesh concepts and container security practices.
  • Experience with cloud or private-cloud infrastructure and distributed systems.

Responsibilities

  • Operate, maintain and troubleshoot large-scale Kubernetes-based production environments.
  • Ensure reliability, scalability, availability and performance of critical services.
  • Investigate complex production issues and perform root-cause analysis.
  • Participate in incident response and drive permanent corrective actions.
  • Automate repetitive operational activities and improve platform reliability.
  • Build and improve monitoring, alerting, logging and observability frameworks.
  • Define and track SLIs, SLOs and operational reliability metrics.
  • Support Kubernetes upgrades, configuration changes and platform improvements.
  • Collaborate with engineering, infrastructure, platform, security and DevOps teams.
  • Perform capacity planning and performance tuning for reliability improvements.
  • Develop and maintain runbooks, automation scripts and troubleshooting docs.
  • Participate in production readiness reviews and ensure apps meet operational standards.

Skills

Kubernetes administration
Kubernetes architecture
Pods
Deployments
StatefulSets
Services
Ingress
ConfigMaps/Secrets
RBAC
Storage
Networking
Linux administration
DNS/TCP/IP
Site Reliability Engineering
Incident management
RCA analysis
Python scripting
Bash scripting
Prometheus
Grafana
Splunk
ELK/OpenSearch
Datadog
Helm
GitOps - Argo CD/Flux
Service mesh concepts
Cloud/private-cloud infra
Distributed systems
Capacity management
Global distributed teams

Tools

Prometheus
Grafana
Splunk
ELK/OpenSearch
Datadog
Argo CD
Flux
Helm
Kubernetes security tools

Job description

Role Overview
  • We are looking for experienced Kubernetes & Site Reliability Engineers to support highly scalable, business-critical production platforms for a global technology customer in Singapore.
  • The role requires strong hands-on expertise in Kubernetes, Linux, production reliability, automation, observability, incident management and troubleshooting of distributed systems
  • Candidates should be comfortable operating large-scale production environments where availability, performance, automation and operational excellence are critical.
Key Responsibilities
  • Operate, maintain and troubleshoot large-scale Kubernetes-based production environments
  • Ensure reliability, scalability, availability and performance of critical services.
  • Investigate complex production issues and perform detailed root-cause analysis.
  • Participate in incident response and drive permanent corrective actions.
  • Automate repetitive operational activities and improve platform reliability.
  • Build and improve monitoring, alerting, logging and observability frameworks.
  • Define and track SLIs, SLOs and operational reliability metrics
  • Support Kubernetes upgrades, configuration changes, patching and platform improvements.
  • Work closely with application engineering, infrastructure, platform, security and DevOps teams.
  • Perform capacity planning, performance tuning and reliability improvements.
  • Develop and maintain operational runbooks, automation scripts and troubleshooting documentation.
  • Participate in production readiness reviews and ensure applications meet operational standards.
Mandatory Skills
  • Strong hands-on experience with
  • Kubernetes administration and troubleshooting
  • Strong understanding of Kubernetes architecture, including:
  • Pods
  • Deployments
  • StatefulSets
  • Services
  • Ingress
  • ConfigMaps / Secrets
  • RBAC
  • Storage
  • Networking
  • Strong
  • Linux systems administration and troubleshooting skills.
  • Good understanding of networking concepts such as DNS, TCP/IP, load balancing and service connectivity.
  • Strong understanding of
  • Site Reliability Engineering principles
  • Experience supporting large-scale, high-availability production systems.
  • Strong incident management and RCA experience.
  • Hands-on scripting/automation experience using
  • Python, Bash/Shell or similar
  • Experience with monitoring and observability tools such as
  • Prometheus, Grafana, Splunk, ELK/OpenSearch, Datadog or equivalent
  • Helm or similar Kubernetes package/deployment management tools.
  • GitOps experience using tools such as Argo CD or Flux.
  • Knowledge of service mesh concepts.
  • Experience with container security and Kubernetes security practices.
  • Experience with cloud or private-cloud infrastructure.
  • Familiarity with distributed systems and microservices architectures.
  • Exposure to performance engineering and capacity management.
  • Experience working in globally distributed engineering environments.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Kubernetes & Site Reliability Engineer (SRE)
Kubernetes & Site Reliability Engineer (SRE)

OPENSOURCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 160,000
Cloud Engineer Kubernetes
Cloud Engineer Kubernetes

OPENSOURCE TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Cloud Engineer- Kubernetes
Cloud Engineer- Kubernetes

OPENSOURCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Platform Engineer / Site Reliability Engineer (Kubernetes)
Platform Engineer / Site Reliability Engineer (Kubernetes)

BOUNTEOUSXACCOLITE SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 170,000
Kubernetes SRE: Reliability, Automation & Observability
Kubernetes SRE: Reliability, Automation & Observability

OPENSOURCE TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Kubernetes SRE & Reliability Engineer
Senior Kubernetes SRE & Reliability Engineer

OPENSOURCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Kubernetes SRE: Reliability Automation Observability
Senior Kubernetes SRE: Reliability Automation Observability

OPENSOURCE TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Senior Kubernetes & Site Reliability Engineer
Senior Kubernetes & Site Reliability Engineer

OPENSOURCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 160,000
DevOps / Site Reliability Engineer
DevOps / Site Reliability Engineer

ACCORD INNOVATIONS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TP-LINK CORPORATION PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000