Senior Site Reliability Engineer

StraitsX

Jakarta Pusat

On-site

IDR 500,000,000 - 900,000,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

StraitsX is seeking a Senior Site Reliability Engineer to own reliability, performance, and cost outcomes across cloud and on‑prem environments. You’ll design deployment strategies for Kubernetes, manage CI/CD and GitOps pipelines, and mentor teammates as a technical reference.

You will diagnose database performance issues, build observability, and drive cross‑team improvements while communicating tradeoffs to stakeholders and leading incident investigations with a metrics‑first approach.

Qualifications

  • At least 4 years of experience in SRE, DevOps, MLOps, or platform engineering at a high-traffic company.

Responsibilities

  • Own the availability, performance, scalability, and security of production systems end-to-end, across cloud (AWS/GCP) and on-premises environments.

Skills

SRE/DevOps
Kubernetes
Linux fundamentals
CI/CD
GitOps
MySQL
PostgreSQL
Datadog
OpenTelemetry
Terraform
AWS

Tools

ArgoCD
Terraform
Datadog

Job description

About The Role

The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. At the Senior Level, you own reliability, performance, and cost outcomes for the systems under your area end-to-end, not just executing well-defined tasks, but deciding between tradeoffs, scoping ambiguous problems, and driving process and system improvements that span teams. You’ll work closely with development, security, and product teams, and mentor other engineers as a technical point of reference for the team.

What You Will Do
  • Own the availability, performance, scalability, and security of production systems end-to-end, across cloud (AWS/GCP) and on-premises environments.
  • Design and evolve Kubernetes deployment strategy for production workloads.
  • Own CI/CD and GitOps pipelines in production (ArgoCD or equivalent) and the Terraform that provisions the infrastructure behind them.
  • Diagnose and resolve database performance issues.
  • Build and maintain observability that surfaces problems before they become incidents.
  • Seek out and implement process and system improvements affecting performance and security, coordinating with multiple stakeholders.
  • Scope and lead medium-to-large infrastructure initiatives: gather requirements, prioritize by business impact, and communicate impact to stakeholders.
  • Negotiate technical tradeoffs with stakeholders to meet business SLAs.
  • Provide technical guidance and mentorship to peers and junior engineers; promote best practices and standards across the team.
  • Maintain documentation and process discipline for the systems and incidents you own.
  • Lead structured incident investigation, isolating server, database, and application layers with a metrics-first approach, including on-call during high-traffic events.
What Are We Looking For
  • At least 4 years of experience in either SRE, DevOps, MLOps, or platform engineering, including senior-level scope at a high-traffic company.
  • Deep expertise in one major cloud provider (preferably AWS), with a proven ability to ramp up on the other quickly.Production experience & expertise withKubernetes &Linux fundamentals
  • CI/CD & GitOps (ArgoCD or other equivalent stacks)
  • Database performance analysis & monitoring (MySQL, Postgres)
  • Observability tooling & standards (Datadog, OpenTelemetry)
  • Infrastructure as Code (Terraform)
  • Strong working English, verbal & written communication. Strong documentation and process discipline.
  • On-call & incident handling experience during high-traffic events.
  • Comfortable negotiating with stakeholders to propose technical compromises that meet business SLAs.
  • Demonstrated growth mindset and proven ability to own ambiguous scope.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

AlloFresh • Jakarta Pusat

On-site
IDR 450,000,000 - 750,000,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
IDR 167,400,000 - 279,000,000
Site Reliability Engineer
Site Reliability Engineer

Pengiklan Anonim • Jakarta Utara

On-site
IDR 446,400,000 - 781,200,000
SRE (Shifting)
SRE (Shifting)

Konnco Studio • Daerah Istimewa Yogyakarta

On-site
IDR 180,000,000 - 280,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
Senior SRE: Lead End-to-End Reliability & Platform
Senior SRE: Lead End-to-End Reliability & Platform

StraitsX • Jakarta Pusat

On-site
IDR 500,000,000 - 900,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AccelByte • Kota Yogyakarta

Hybrid
IDR 420,000,000 - 700,000,000
Senior/Staff DevOps Engineer / SRE
Senior/Staff DevOps Engineer / SRE

Ajaib • Jakarta Pusat

On-site
IDR 300,000,000 - 600,000,000
Infrastructure Site Reliability Engineer
Infrastructure Site Reliability Engineer

PT Brokeret Fintech Solutions • Tangerang

On-site
IDR 250,000,000 - 480,000,000
DevOps Engineer
DevOps Engineer

SkyRocket Infosystem • Lembang

On-site
IDR 300,000,000 - 600,000,000
SOC 2/ISO 27001 experience
Service mesh
FinOps experience