Senior Site Reliability Engineer

StraitsX Group

Jakarta Pusat

On-site

IDR 400,000,000 - 700,000,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

StraitsX Group is seeking a Senior Site Reliability Engineer to own reliability, performance, and cost outcomes for production systems. You will work across cloud and on‑prem environments, design Kubernetes deployment strategies, and lead CI/CD pipelines and Terraform provisioning.

Mentorship and incident leadership are key components of the role. You will collaborate with development, security, and product teams, drive observability improvements, and negotiate technical tradeoffs to meet SLAs.

Qualifications

  • 4+ years in SRE/DevOps/Platform engineering at a high-traffic company.
  • Deep expertise in one cloud (AWS preferred) with Kubernetes and Linux fundamentals.
  • Experience with CI/CD and GitOps (ArgoCD or equivalent).
  • Database performance analysis and monitoring (MySQL, Postgres).
  • Observability tooling (Datadog, OpenTelemetry) and IaC (Terraform).
  • Strong English communication; on-call and incident handling experience.

Responsibilities

  • Own availability, performance, scalability, and security end-to-end across cloud and on-prem.
  • Design and evolve Kubernetes deployment strategy for production workloads.
  • Own CI/CD and GitOps pipelines (ArgoCD) and the Terraform provisioning.
  • Diagnose and resolve database performance issues.
  • Build and maintain observability to surface problems before incidents.
  • Lead infrastructure initiatives across teams; prioritize by business impact.
  • Negotiate tradeoffs to meet business SLAs.
  • Provide mentorship and promote best practices.
  • Maintain docs and process discipline; conduct structured incident investigations and on-call during high-traffic events.

Skills

AWS
Kubernetes
Linux
CI/CD
GitOps
Datadog
OpenTelemetry
On-call

Tools

ArgoCD
Terraform

Job description

About The Role

The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. At the Senior Level, you own reliability, performance, and cost outcomes for the systems under your area end-to-end, not just executing well-defined tasks, but deciding between tradeoffs, scoping ambiguous problems, and driving process and system improvements that span teams. You'll work closely with development, security, and product teams, and mentor other engineers as a technical point of reference for the team.

What You Will Do
  • Own the availability, performance, scalability, and security of production systems end-to-end, across cloud (AWS/GCP) and on-premises environments.
  • Design and evolve Kubernetes deployment strategy for production workloads.
  • Own CI/CD and GitOps pipelines in production (ArgoCD or equivalent) and the Terraform that provisions the infrastructure behind them.
  • Diagnose and resolve database performance issues.
  • Build and maintain observability that surfaces problems before they become incidents.
  • Seek out and implement process and system improvements affecting performance and security, coordinating with multiple stakeholders.
  • Scope and lead medium-to-large infrastructure initiatives: gather requirements, prioritize by business impact, and communicate impact to stakeholders.
  • Negotiate technical tradeoffs with stakeholders to meet business SLAs.
  • Provide technical guidance and mentorship to peers and junior engineers; promote best practices and standards across the team.
  • Maintain documentation and process discipline for the systems and incidents you own.
  • Lead structured incident investigation, isolating server, database, and application layers with a metrics-first approach, including on-call during high-traffic events.
What Are We Looking For
  • At least 4 years of experience in either SRE, DevOps, MLOps, or platform engineering, including senior-level scope at a high-traffic company.
  • Deep expertise in one major cloud provider (preferably AWS), with a proven ability to ramp up on the other quickly. Production experience & expertise with Kubernetes & Linux fundamentals
  • CI/CD & GitOps (ArgoCD or other equivalent stacks)
  • Database performance analysis & monitoring (MySQL, Postgres)
  • Observability tooling & standards (Datadog, OpenTelemetry)
  • Infrastructure as Code (Terraform)
  • Strong working English, verbal & written communication. Strong documentation and process discipline.
  • On-call & incident handling experience during high-traffic events.
  • Comfortable negotiating with stakeholders to propose technical compromises that meet business SLAs.
  • Demonstrated growth mindset and proven ability to own ambiguous scope.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer New Jakarta, Jakarta, Indonesia
Senior Site Reliability Engineer New Jakarta, Jakarta, Indonesia

Greenhouse Software, Inc. • Jakarta Pusat

On-site
IDR 400,000,000 - 800,000,000
Site Reliability Engineer
Site Reliability Engineer

AlloFresh • Jakarta Pusat

On-site
IDR 450,000,000 - 750,000,000
Site Reliability Engineer
Site Reliability Engineer

Pengiklan Anonim • Jakarta Utara

On-site
IDR 446,400,000 - 781,200,000
Executive - Site Reliability Engineer
Executive - Site Reliability Engineer

Macquarie Group • Indonesia

On-site
IDR 350,000,000 - 550,000,000
Senior SRE: End-to-End Reliability & Platform Lead
Senior SRE: End-to-End Reliability & Platform Lead

Greenhouse Software, Inc. • Jakarta Pusat

On-site
IDR 400,000,000 - 800,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AccelByte • Kota Yogyakarta

On-site
IDR 420,000,000 - 700,000,000
Senior/Staff DevOps Engineer / SRE
Senior/Staff DevOps Engineer / SRE

Ajaib • Jakarta Pusat

On-site
IDR 300,000,000 - 600,000,000
Infrastructure Site Reliability Engineer
Infrastructure Site Reliability Engineer

PT Brokeret Fintech Solutions • Tangerang

On-site
IDR 250,000,000 - 480,000,000
Site Reliability Engineer - LInE (Remote)
Site Reliability Engineer - LInE (Remote)

Quik Hire Staffing • Indonesia

Remote
IDR 260,000,000 - 520,000,000
Senior Site Reliability Engineer - Lead Reliability & Scale
Senior Site Reliability Engineer - Lead Reliability & Scale

StraitsX Group • Jakarta Pusat

On-site
IDR 400,000,000 - 700,000,000