Senior Site Reliability Engineer

StraitsX

Indonesia

On-site

IDR 300,000,000 - 540,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

StraitsX is seeking a Senior Site Reliability Engineer to own reliability, performance, and cost outcomes across cloud (AWS/GCP) and on-premises environments. You will design Kubernetes deployment strategies, manage CI/CD pipelines with ArgoCD, and maintain Terraform-backed infrastructure while coaching peers and leading major infrastructure initiatives.

You will diagnose database performance issues, build robust observability with Datadog and OpenTelemetry, and drive process improvements across

Qualifications

  • 4+ years in SRE/DevOps/ML Ops or platform engineering with senior scope.
  • Deep expertise in one major cloud (preferably AWS) with ability to ramp up others.
  • Production experience with Kubernetes and Linux fundamentals.
  • CI/CD and GitOps (ArgoCD or equivalents).
  • Database performance analysis & monitoring (MySQL, PostgreSQL).
  • Observability tooling (Datadog, OpenTelemetry).
  • Infrastructure as Code (Terraform).
  • Strong written and verbal English; solid documentation and process discipline.
  • On-call and incident handling during high-traffic events.
  • Ability to negotiate tradeoffs to meet SLAs and own ambiguous scope.

Responsibilities

  • Own availability, performance, scalability, and security of production systems end-to-end.
  • Design and evolve Kubernetes deployment strategy for production workloads.
  • Own CI/CD and GitOps pipelines and the Terraform that provisions the infra.
  • Diagnose and resolve database performance issues.
  • Build observability to surface problems before incidents.
  • Drive process and system improvements across teams.
  • Lead medium-to-large infrastructure initiatives with stakeholder communication.
  • Negotiate technical tradeoffs to meet business SLAs.
  • Provide mentorship and promote best practices across the team.
  • Maintain documentation and incident response discipline.
  • Lead structured incident investigations with metrics-first approach; on-call during high traffic.

Skills

SRE
DevOps
MLOps
Kubernetes
Linux fundamentals
CI/CD
GitOps
AWS
Datadog
OpenTelemetry
Terraform
MySQL
PostgreSQL
ArgoCD
English communication
On-call

Tools

ArgoCD
Terraform
Datadog
OpenTelemetry

Job description

The Site Reliability Engineering (SRE) team architects, builds, and maintains the rock-solid infrastructure that applications rely on. At the Senior Level, you own reliability, performance, and cost outcomes for the systems under your area end-to-end, not just executing well-defined tasks, but deciding between tradeoffs, scoping ambiguous problems, and driving process and system improvements that span teams. You'll work closely with development, security, and product teams, and mentor other engineers as a technical point of reference for the team.

What You Will Do
  • Own the availability, performance, scalability, and security of production systems end-to-end, across cloud (AWS/GCP) and on-premises environments.
  • Design and evolve Kubernetes deployment strategy for production workloads.
  • Own CI/CD and GitOps pipelines in production (ArgoCD or equivalent) and the Terraform that provisions the infrastructure behind them.
  • Diagnose and resolve database performance issues.
  • Build and maintain observability that surfaces problems before they become incidents.
  • Seek out and implement process and system improvements affecting performance and security, coordinating with multiple stakeholders.
  • Scope and lead medium-to-large infrastructure initiatives: gather requirements, prioritize by business impact, and communicate impact to stakeholders.
  • Negotiate technical tradeoffs with stakeholders to meet business SLAs.
  • Provide technical guidance and mentorship to peers and junior engineers; promote best practices and standards across the team.
  • Maintain documentation and process discipline for the systems and incidents you own.
  • Lead structured incident investigation, isolating server, database, and application layers with a metrics-first approach, including on-call during high-traffic events.
What Are We Looking For
  • At least 4 years of experience in either SRE, DevOps, MLOps, or platform engineering, including senior-level scope at a high-traffic company.
  • Deep expertise in one major cloud provider (preferably AWS), with a proven ability to ramp up on the other quickly.Production experience & expertise with Kubernetes & Linux fundamentals
  • CI/CD & GitOps (ArgoCD or other equivalent stacks)
  • Database performance analysis & monitoring (MySQL, Postgres)
  • Observability tooling & standards (Datadog, OpenTelemetry)
  • Infrastructure as Code (Terraform)
  • Strong working English, verbal & written communication. Strong documentation and process discipline.
  • On-call & incident handling experience during high-traffic events.
  • Comfortable negotiating with stakeholders to propose technical compromises that meet business SLAs.
  • Demonstrated growth mindset and proven ability to own ambiguous scope.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Quiet Capital • Jakarta Pusat

On-site
IDR 350,000,000 - 700,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

StraitsX • Jakarta Pusat

On-site
IDR 600,000,000 - 840,000,000
Site Reliability Engineer
Site Reliability Engineer

StraitsX • Daerah Khusus Ibukota Jakarta

On-site
Senior Site Reliability Engineer New Jakarta, Jakarta, Indonesia
Senior Site Reliability Engineer New Jakarta, Jakarta, Indonesia

StraitsX Group • Jakarta Pusat

On-site
IDR 550,000,000 - 750,000,000
Site Reliability Engineer
Site Reliability Engineer

PARTECH PARTNERS • Daerah Khusus Ibukota Jakarta

On-site
IDR 272,380,000 - 453,968,000
Site Reliability Engineer
Site Reliability Engineer

AlloFresh • Jakarta Pusat

On-site
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BookCabin • Jakarta Pusat

On-site
IDR 250,000,000 - 420,000,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
IT Site Reliability Engineer (SRE)
IT Site Reliability Engineer (SRE)

PT ABISHAR TECHNOLOGIES INDONESIA • Jakarta Barat

On-site
IDR 36,000,000 - 60,000,000
Senior SRE: End-to-End Reliability & Platform Lead
Senior SRE: End-to-End Reliability & Platform Lead

StraitsX Group • Jakarta Pusat

On-site
IDR 550,000,000 - 750,000,000