Site Reliability Engineer

AlloFresh

Jakarta Pusat

On-site

IDR 450,000,000 - 750,000,000

Full time

9 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

AlloFresh is seeking a Site Reliability Engineer to ensure reliability, scalability, and performance of production systems in a cloud-native environment. You will implement SRE practices, automate infrastructure, and lead incident response while partnering with development teams to embed reliability across the software lifecycle.

You will manage Kubernetes clusters, maintain GitOps workflows with ArgoCD, design CI/CD pipelines with Jenkins, and contribute to runbooks, postmortems, and

Qualifications

  • 3+ years in Site Reliability Engineering (SRE), DevOps, or infrastructure engineering.
  • Solid understanding of SRE principles, including SLIs, SLOs, and error budget frameworks.
  • Experience operating and managing systems in cloud-native environments (e.g., AWS, GCP, Azure).
  • Strong proficiency in containerization and orchestration technologies, particularly Docker and Kubernetes.
  • Experience with CI/CD pipelines, automation, and deployment strategies (e.g., Jenkins, GitLab CI, GitHub Actions).
  • Familiarity with GitOps practices and tools such as ArgoCD or FluxCD.
  • Hands-on experience with observability and monitoring tools (e.g., Prometheus, Grafana, logging and alerting systems).
  • Proficiency in scripting and automation (e.g., Bash, Python, Groovy) and Infrastructure as Code (e.g., Terraform, Ansible).
  • Strong foundation in Linux/Unix systems and Git-based version control workflows.

Responsibilities

  • Monitor system reliability and define, implement, and track SLIs, SLOs, and error budgets for production services.
  • Lead incident response activities, including troubleshooting, root cause analysis, and resolution.
  • Automate infrastructure processes and reduce operational toil through scripting and tooling enhancements.
  • Manage and optimize Kubernetes clusters, including resource configurations, manifests, and kustomize overlays.
  • Maintain and monitor GitOps workflows using ArgoCD to ensure consistent and reliable application deployments.
  • Design, build, and maintain CI/CD pipelines using Jenkins Job DSL (Groovy).
  • Collaborate with development teams to improve system reliability, scalability, and deployment readiness.
  • Maintain and upgrade platform tools such as Jenkins, ArgoCD, and container registries.
  • Develop and maintain runbooks, postmortems, and internal technical documentation.

Skills

SLI/SLO practices
Kubernetes
CI/CD pipelines
GitOps
Observability
Scripting
IaC
Linux

Job description

As a Site Reliability Engineer, you will be responsible for ensuring the reliability, scalability, and performance of production systems. You will apply engineering principles to operations, focusing on automation, observability, and continuous improvement to reduce manual effort and enhance system resilience. You will work closely with development teams to build and operate highly reliable services, embedding reliability into the full software lifecycle.

Main Responsibilities
  • Monitor system reliability and define, implement, and track SLIs, SLOs, and error budgets for production services
  • Lead incident response activities, including troubleshooting, root cause analysis, and resolution
  • Automate infrastructure processes and reduce operational toil through scripting and tooling enhancements
  • Manage and optimize Kubernetes clusters, including resource configurations, manifests, and kustomize overlays
  • Maintain and monitor GitOps workflows using ArgoCD to ensure consistent and reliable application deployments
  • Design, build, and maintain CI/CD pipelines using Jenkins Job DSL (Groovy)
  • Collaborate with development teams to improve system reliability, scalability, and deployment readiness
  • Maintain and upgrade platform tools such as Jenkins, ArgoCD, and container registries
  • Develop and maintain runbooks, postmortems, and internal technical documentation
Requirements
  • 3+ years of experience in Site Reliability Engineering (SRE), DevOps, or infrastructure engineering
  • Solid understanding of SRE principles, including SLIs, SLOs, and error budget frameworks
  • Experience operating and managing systems in cloud-native environments (e.g., AWS, GCP, Azure)
  • Strong proficiency in containerization and orchestration technologies, particularly Docker and Kubernetes
  • Experience with CI/CD pipelines, automation, and deployment strategies (e.g., Jenkins, GitLab CI, GitHub Actions)
  • Familiarity with GitOps practices and tools such as ArgoCD or FluxCD
  • Hands-on experience with observability and monitoring tools (e.g., Prometheus, Grafana, logging and alerting systems)
  • Proficiency in scripting and automation (e.g., Bash, Python, Groovy) and Infrastructure as Code (e.g., Terraform, Ansible)
  • Strong foundation in Linux/Unix systems and Git-based version control workflows
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

BookCabin • Jakarta Pusat

On-site
IDR 250,000,000 - 420,000,000
Site Reliability Engineer
Site Reliability Engineer

Pengiklan Anonim • Jakarta Utara

On-site
IDR 446,400,000 - 781,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

StraitsX • Jakarta Pusat

On-site
IDR 600,000,000 - 840,000,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
Senior Site Reliability Engineer New Jakarta, Jakarta, Indonesia
Senior Site Reliability Engineer New Jakarta, Jakarta, Indonesia

StraitsX Group • Jakarta Pusat

On-site
IDR 550,000,000 - 750,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AccelByte • Sleman

On-site
SRE: Architect Reliable, Scalable Infrastructure
SRE: Architect Reliable, Scalable Infrastructure

StraitsX Group • Jakarta Pusat

On-site
IDR 200,000,000 - 300,000,000
Site Reliability Engineer New Jakarta, Jakarta, Indonesia
Site Reliability Engineer New Jakarta, Jakarta, Indonesia

StraitsX Group • Jakarta Pusat

On-site
IDR 200,000,000 - 300,000,000
Site Reliability Engineer (Junior)
Site Reliability Engineer (Junior)

CloudMile • Jakarta Pusat

On-site
IDR 178,560,000 - 245,520,000
Site Reliability Engineer / DevOps
Site Reliability Engineer / DevOps

Catalyst Tech • Daerah Khusus Ibukota Jakarta

On-site
IDR 720,720,000 - 1,081,082,000