Senior Site Reliability (DevOps) Engineer

Rakuten Asia Pte Ltd

Singapore

On-site

SGD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Rakuten Asia Pte Ltd, located in Singapore, is seeking a Senior Site Reliability (DevOps) Engineer to bridge engineering and operations. This role is crucial in defining SRE strategy, improving incident management processes, and driving automation initiatives. The ideal candidate will possess over 8 years of experience in software engineering, with at least 3 years in a management role, and expertise in cloud platforms.

This position emphasizes creating reliable systems for millions of Rakuten's customers globally while collaborating effectively across teams. Comprehensive knowledge of containerization technologies and observability practices is essential for success in this position.

Qualifications

  • 8+ years of experience in software engineering, DevOps, or site reliability engineering, with at least 3 years in a management role.
  • Proven track record of leading high-performing teams in a distributed environment.
  • Deep expertise in cloud platforms including GCP, AWS, or Azure.
  • Strong knowledge of containerization technologies and Infrastructure as Code.
  • Hands-on experience with observability tools and defining SLOs/SLIs.

Responsibilities

  • Define and drive SRE strategy, focusing on reliability targets.
  • Establish incident-management processes to minimize MTTR.
  • Collaborate with teams to embed reliability in the software lifecycle.
  • Design observability solutions to provide insights into system health.
  • Drive automation initiatives for improved reliability and efficiency.
  • Partner with Architecture teams to support scalability and costs.
  • Manage capacity planning for marketing platforms.
  • Report reliability metrics and translate insights into business impact.

Skills

Software engineering
DevOps
Site reliability engineering
Cloud platforms (GCP preferred)
Containerization (Kubernetes, Docker)
Infrastructure as Code (Terraform, Ansible)
Observability tools (Prometheus, Grafana)
CI/CD pipelines
Programming/scripting (Python, Go, Java)
Incident management

Job description

Situated in the heart of Singapore's Central Business District, Rakuten Asia Pte. Ltd. is Rakuten's Asia Regional headquarters. Established in August 2012 as part of Rakuten's global expansion strategy, Rakuten Asia comprises various businesses that provide essential value-added services to Rakuten's global ecosystem. Through advertisement product development, product strategy, and data management, among others, Rakuten Asia is strengthening Rakuten Group's core competencies to lead in an increasingly digitalized world.

Rakuten Group, Inc. is a global leader in internet services that empower individuals, communities, businesses, and society. Founded in Tokyo in 1997 as an online marketplace, Rakuten has expanded to offer services in e‑commerce, fintech, digital content, and communications to approximately 1.7 billion members worldwide. The Rakuten Group has nearly 32,000 employees and operations in 30 countries and regions.

The Marketing Cloud Platform Department (MCPD) drives Rakuten's marketing product strategy, executes product development, and ensures successful implementation. We empower Rakuten's internal marketing teams by creating engaging, respectful, and cost‑efficient marketing platforms that prioritize our customers. Leveraging the Rakuten Ecosystem, we offer comprehensive marketing solutions, including campaign management, multichannel communication, and personalization.

Senior Site Reliability (DevOps) Engineer

This role bridges engineering and operations, requiring both strong technical expertise and people‑management skills to build and maintain highly available systems that serve millions of Rakuten's customers globally.

Responsibilities
  • Define and drive SRE strategy, including SLO/SLI frameworks, error budgets, and reliability targets aligned with business objectives and customer expectations.
  • Establish and improve incident‑management processes, including on‑call rotations, escalation procedures, and blameless post‑mortem practices to minimize MTTR and prevent recurring issues.
  • Collaborate with development teams to embed reliability practices into the software development lifecycle, advocating for design reviews, chaos engineering, and production readiness reviews.
  • Design and implement comprehensive observability solutions (monitoring, logging, tracing, alerting) to provide actionable insights into system health and performance.
  • Drive automation initiatives to reduce toil, improve deployment reliability, and enable self‑service capabilities for engineering teams.
  • Partner with Architecture and Platform teams to ensure infrastructure decisions support scalability, fault tolerance, and cost optimization goals.
  • Manage capacity planning and performance optimization for critical marketing platforms handling high‑volume campaign executions and real‑time personalization.
  • Report on reliability metrics, incident trends, and operational health to leadership, translating technical insights into business impact assessments.
Required Qualifications
  • 8+ years of experience in software engineering, DevOps, or site reliability engineering, with at least 3 years in a people‑management role.
  • Proven track record of building and leading high‑performing SRE or platform engineering teams in a distributed, multi‑timezone environment.
  • Deep expertise in cloud platforms (GCP preferred, AWS/Azure acceptable) including compute, networking, storage, and managed services.
  • Strong knowledge of containerization and orchestration technologies (Kubernetes, Docker) and Infrastructure as Code (Terraform, Ansible).
  • Hands‑on experience with observability tools and practices (Prometheus, Grafana, Datadog, ELK Stack, or similar) and defining meaningful SLOs/SLIs.
  • Experience with CI/CD pipelines, deployment strategies (blue‑green, canary), and release engineering best practices.
  • Strong programming/scripting skills in languages such as Python, Go, or Java for automation and tooling development.
  • Excellent communication skills with the ability to collaborate effectively across engineering, product, and business stakeholders.
  • Strong incident‑management experience with demonstrated ability to lead high‑pressure situations calmly and effectively.
Nice to Have
  • Experience with big data technologies (Hadoop, Spark, Kafka) and data pipeline reliability.
  • Familiarity with marketing technology platforms, email delivery systems, or customer data platforms.
  • Knowledge of database administration and optimization (PostgreSQL, MySQL, Redis, Couchbase).
  • Experience with chaos engineering practices and tools (Chaos Monkey, Litmus, Gremlin).
  • Certifications such as Google Cloud Professional Cloud Architect, AWS Solutions Architect, or Kubernetes Administrator (CKA).
  • Japanese language proficiency is a plus for collaboration with Japan‑based teams.

Rakuten is an equal opportunities employer and welcomes applications regardless of sex, marital status, ethnic origin, sexual orientation, religious belief, or age.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE & Reliability Engineer - Lead & Automate
Senior SRE & Reliability Engineer - Lead & Automate

Rakuten Asia Pte Ltd • Singapore

On-site
SGD 120,000 - 160,000
Associate Engineer/Engineer, SRE
Associate Engineer/Engineer, SRE

Rakuten Viki • Singapore

On-site
SGD 90,000 - 130,000
Cloud Specialist
Cloud Specialist

VIKI PRIVATE LIMITED • Singapore

On-site
SGD 120,000 - 180,000
Senior Project Manager
Senior Project Manager

Rakuten Asia Pte Ltd • Singapore

On-site
SGD 90,000 - 120,000
Senior Software Engineer
Senior Software Engineer

Rakuten Asia Pte Ltd • Singapore

On-site
SGD 70,000 - 90,000
Engineering Manager, Data Engineering
Engineering Manager, Data Engineering

Rakuten Viki • Singapore

On-site
SGD 120,000 - 160,000
Vice Senior Manager
Vice Senior Manager

Rakuten Kobo Inc. • Singapore

On-site
SGD 150,000 - 210,000
Vice Senior Manager
Vice Senior Manager

Rakuten Asia Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
Senior Corporate Planning Manager
Senior Corporate Planning Manager

Rakuten Kobo Inc. • Singapore

On-site
SGD 180,000 - 240,000
Senior Software Engineer (APS)
Senior Software Engineer (APS)

Rakuten Kobo Inc. • Singapore

On-site
SGD 90,000 - 120,000