Site Reliability Engineer

DANA Indonesia

Jakarta Pusat

On-site

IDR 156,240,000 - 267,840,000

Full time

10 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

DANA Indonesia in Jakarta is building a robust Business Automation platform and KYB services. This role ensures reliable, resilient systems while enabling engineering leadership to focus on architecture, design reviews and quality assurance.

You will operate Kubernetes workloads, implement monitoring and incident response, write runbooks, and automate recurring tasks while leveraging IaC and cloud tooling to maintain uptime.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, Information Systems or related field or equivalent practical experience.
  • 1–2 years of experience in systems administration, DevOps, infrastructure or IT operations.
  • Hands-on Linux administration, networking, storage, performance troubleshooting, and CLI operations.

Responsibilities

  • Operate and maintain Automation platform day-to-day, including Kubernetes workloads, ingress, load balancing and CI/CD pipelines.
  • Build and tune monitoring and alerting; ensure ownership and clear response paths for alerts.
  • Write and maintain runbooks and operational docs for services with little documentation.
  • Automate recurring operational work to reduce manual effort.
  • Support infrastructure provisioning via IaC and review changes with Engineering Manager.
  • Assist root cause analysis using logs, metrics and system behavior.
  • Support backup, restore and disaster recovery exercises and verify recovery procedures.
  • Gradually take ownership of defined operational areas as familiarity grows.

Skills

Linux admin
Networking basics
Scripting: Python/Go/Bash
Terraform/Ansible
Docker & Kubernetes
CI/CD (GitHub Actions/GitLab)
Observability tools
Incident response
Runbooks & docs

Education

Bachelor’s degree in CS/Engineering or related field

Tools

Docker
Kubernetes
Terraform
Ansible
Git
GitHub Actions
GitLab CI
Grafana
OpenSearch/ELK

Job description

Own the day-to-day reliability and operations of the Business Automation platform and KYB services, ensuring stable, resilient systems and reducing dependency on individual team members. This role strengthens operational ownership across the team while enabling engineering leadership to focus on architecture, design reviews, and quality assurance.

You'll be working on:

  • Operate and maintain the Automation platform running reliably day to day, including Kubernetes workloads, ingress and load balancing, secrets management, and CI/CD pipelines.
  • Build and tune monitoring and alerting so failures are detected automatically before someone has to notice them manually, and ensure every alert has a clear owner and response path.
  • Write and maintain runbooks and operational documentation, starting with services that currently have little or no documentation.
  • Automate recurring operational work, reducing manual work and repetitive load on the team.
  • Support infrastructure provisioning and configuration through Infrastructure as Code, reviewing changes with the Engineering Manager.
  • Assist root cause analysis after incidents by gathering and reviewing evidence from logs, metrics, and system behaviour.
  • Support backup, restore and disaster recovery exercises, and make sure recovery procedures actually work when needed.
  • Gradually take independent ownership of defined operational areas as you build deeper familiarity with the platform.

Qualifications:

  • Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience.
  • 1–2 years of experience in systems administration, DevOps, infrastructure, or IT operations. Strong internship or project experience may also be considered.
  • Hands-on knowledge of Linux administration, including permissions, networking, storage, performance troubleshooting, and command-line operations.
  • Good understanding of networking fundamentals, including TCP/IP, DNS, routing, load balancing, and firewalls, with the ability to troubleshoot connectivity and latency issues.
  • Scripting skills in Python, Go, or Bash to automate operational and infrastructure tasks.
  • Familiarity with virtualization, Docker, and Kubernetes; experience operating containerized workloads in production is preferred.
  • Exposure to Infrastructure as Code, particularly Terraform or Ansible, and version control using Git.
  • Familiarity with CI/CD pipelines such as GitHub Actions or GitLab CI, including automated testing, builds, and deployment workflows.
  • Exposure to cloud infrastructure, ideally Alibaba Cloud or Azure; experience with AWS or GCP is also relevant.
  • Familiarity with observability and monitoring tools such as Grafana, OpenSearch, or the ELK stack.
  • Strong troubleshooting mindset with a methodical approach to diagnosing production and infrastructure issues.
  • Clear written and verbal communication, including the ability to document systems, share knowledge, collaborate with developers and product teams, and communicate calmly during incidents.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

amIT Global Solutions Sdn Bhd • Indonesia

On-site
IDR 150,000,000 - 270,000,000
Site Reliability Engineer
Site Reliability Engineer

AlloFresh • Jakarta Pusat

On-site
IDR 450,000,000 - 750,000,000
SRE: Architect of Reliable, Automated Platform
SRE: Architect of Reliable, Automated Platform

amIT Global Solutions Sdn Bhd • Indonesia

On-site
IDR 150,000,000 - 270,000,000
System Operation (on-site)
System Operation (on-site)

Vascomm • Sidoarjo

On-site
IDR 180,000,000 - 240,000,000
Site Reliability Engineer
Site Reliability Engineer

Pengiklan Anonim • Jakarta Utara

On-site
IDR 446,400,000 - 781,200,000
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

PT Media Indonusa (Jakarta) • Jakarta Utara

On-site
IDR 420,000,000 - 540,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AccelByte • Kota Yogyakarta

Hybrid
IDR 420,000,000 - 700,000,000
Head of SysOps
Head of SysOps

PT Media Indonusa (Jakarta) • Jakarta Utara

On-site
IDR 350,000,000 - 550,000,000
Site Reliability Engineer
Site Reliability Engineer

CIMB Niaga • Tangerang Selatan

On-site
IDR 350,000,000 - 550,000,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
IDR 167,400,000 - 279,000,000