Site Reliability Engineer (Junior)

CloudMile

Jakarta Pusat

On-site

IDR 178,560,000 - 245,520,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CloudMile is seeking a skilled SRE Engineer to design, implement, and maintain robust infrastructure on GCP and AWS. You will join a team delivering managed services and ensuring platform reliability for customers.

You will work on monitoring, automation, IaC, CI/CD, and incident response, collaborating with customer success to address complex issues. English and Bahasa Indonesia communication is essential.

Qualifications

  • Fresh graduates or professionals with up to 1–2 years of experience as SRE/DevOps.
  • Experience using GCP, AWS, Alibaba, or Azure.
  • Strong skills in monitoring, logging, and alerting.
  • Proficient in IaC tools and scripting.
  • Experience with containers and orchestration.
  • CI/CD and automation mindset.
  • English and Bahasa Indonesia communication.

Responsibilities

  • Design, build, and maintain scalable platform components.
  • Define, monitor, and report on SLIs/SLOs.
  • Eliminate single points of failure across the infra.
  • Participate in a rotating on-call schedule for after-hours support.
  • Improve monitoring, logging, and alerting for visibility.
  • Automate operational tasks related to platform management.
  • Develop IaC and CI/CD pipelines for reliable infrastructure.
  • Collaborate with Customer Success to resolve complex issues.
  • Contribute to architecture focusing on resilience and multi-tenancy.
  • Perform capacity planning for growth.

Skills

Cloud platforms
Monitoring & alerts
Scripting
Containers
Kubernetes
CI/CD
On-call experience
English Indonesian

Tools

Prometheus
Grafana
ELK Stack
Datadog
Terraform
CloudFormation
Pulumi
Docker
Kubernetes
Swarm

Job description

We are seeking a highly skilled SRE Engineer to join our team and play a critical role in delivering exceptional managed services to our clients. As a key member of our engineering team, you will be responsible for designing, implementing, and maintaining robust and scalable infrastructure solutions on Google Cloud Platform (GCP) and Amazon Web Services (AWS).

Key Responsibilities:
  • Design, build, and maintain highly available, scalable, and performant platform components and shared services.
  • Define, monitor, and report on key Service Level Indicators (SLIs) and Service Level Objectives (SLOs) relevant to platform health and customer experience.
  • Identify and eliminate single points of failure across the infrastructure.
  • Participate in a rotating on-call schedule to provide after-hours support and incident response as needed.
  • Implement and improve monitoring, logging, and alerting systems to gain deep visibility into platform health, resource utilization, and potential issues including those triggered by customer activity.
  • Automate repetitive operational tasks ("toil") related to platform management, provisioning, scaling, and healing.
  • Develop and maintain Infrastructure as Code (IaC) and CI/CD pipeline to manage the platform infrastructure consistently and reliably.
  • Participate in incident response, troubleshooting, and resolution efforts for platform issues.
  • Collaborate with Customer Success teams to diagnose and resolve complex platform issues that may be related to customer-specific configurations or usage.
  • Contribute to the architectural design and evolution of the platform, focusing on resilience, multi-tenancy best practices, and supportability under varying customer loads.
  • Perform capacity planning to ensure the platform can handle anticipated customer growth and usage patterns.
Qualifications:
  • Fresh graduates or professionals with up to 1–2 years of experience as Site Reliability Engineer, DevOps Engineer, or similar role supporting production systems.
  • Experience working with cloud platforms (GCP, AWS, Alibaba, Azure).
  • Strong understanding of monitoring, logging, and alerting principles and tools (Prometheus, Grafana, ELK Stack, Datadog).
  • Proficiency in Infrastructure as Code (IaC) tools (Terraform, CloudFormation, Pulumi).
  • Solid scripting and automation skills (Python, Go, Bash).
  • Experience with containerization and orchestration technologies (Docker, Swarm, Kubernetes).
  • Familiarity with CI/CD pipelines and practices.
  • Understanding of networking fundamentals, databases, and distributed systems.
  • Experience participating in on-call rotations.
  • Excellent problem-solving and troubleshooting skills.
  • Strong communication and collaboration skills.
  • A passion for automation and continuous improvement.
  • A proactive approach in problem identification and resolution - don’t wait around, grab & fix it.
  • A learning attitude.
  • Excellent communication skills in English and Bahasa Indonesia.
Preferred Qualifications:
  • Certifications in GCP or AWS.
  • Experience working directly with customer-facing teams.
  • Experience defining and tracking customer-facing SLOs.
  • Experience providing self-service tooling or observability insights to customers.
  • Experience with cloud cost optimization strategies.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer New Jakarta, Jakarta, Indonesia
Senior Site Reliability Engineer New Jakarta, Jakarta, Indonesia

StraitsX Group • Jakarta Pusat

On-site
IDR 550,000,000 - 750,000,000
Site Reliability Engineer New Jakarta, Jakarta, Indonesia
Site Reliability Engineer New Jakarta, Jakarta, Indonesia

StraitsX Group • Jakarta Pusat

On-site
IDR 200,000,000 - 300,000,000
Lead DevOps Engineer
Lead DevOps Engineer

Enterprise Digital Technology Services Edts • Jakarta Utara

On-site
IDR 390,600,000 - 725,400,000
On-site work at Wisma 46, Jakarta Pus
Competitive compensation
Collaborative environment
Site Reliability Engineer - Jakarta Department Tech Employment Type
Site Reliability Engineer - Jakarta Department Tech Employment Type

Stockbit • Daerah Khusus Ibukota Jakarta

On-site
Site Reliability Engineer – Platform Engineer
Site Reliability Engineer – Platform Engineer

Blibli • Jakarta Timur

On-site
IDR 629,383,000 - 989,031,000
Site Reliability Engineer
Site Reliability Engineer

PARTECH PARTNERS • Daerah Khusus Ibukota Jakarta

On-site
IDR 272,380,000 - 453,968,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

StraitsX • Indonesia

On-site
IDR 300,000,000 - 540,000,000
DevOps Engineer
DevOps Engineer

Pensieve • Indonesia

On-site
IDR 180,000,000 - 300,000,000
Cloud SRE: Build Resilient, Scalable Infra
Cloud SRE: Build Resilient, Scalable Infra

CloudMile • Jakarta Pusat

On-site
IDR 178,560,000 - 245,520,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Quiet Capital • Jakarta Pusat

On-site
IDR 350,000,000 - 700,000,000