Site Reliability Engineer III (Data Platform)

Guidewire Software

Kuala Lumpur

On-site

MYR 80,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Guidewire Software is seeking a Site Reliability Engineer – III (Data Platform) in Kuala Lumpur, Malaysia. The role involves improving incident response, optimizing data processing and analytics platforms on AWS, and ensuring the reliability of critical services.

Candidates should have 4–6 years of experience in SRE or DevOps, with proficiency in managing big data services like Kafka and Spark. This position offers a collaborative environment focused on operational excellence and AI/cloud adoption.

Qualifications

  • 4–6 years of experience in Site Reliability Engineering, DevOps, or Production Engineering.
  • Ability to solve infrastructure problems using software engineering.

Responsibilities

  • Collaborate with senior SREs to enhance incident response.
  • Support the maintenance and enhancement of production environments.
  • Manage and optimize big data and streaming platforms on AWS.
  • Participate in on-call rotations to maintain high availability of services.
  • Resolve production issues through cross-team collaboration.

Skills

Experience deploying and operating services on AWS or Azure
Expertise with data and streaming platforms like Kafka, Hadoop, Spark, or Hive
Proficiency in Go, Python, or Bash for automation
Familiarity with CI/CD tools such as TeamCity, GitHub Actions, or Jenkins
Working knowledge of Kubernetes and Docker

Education

BS/MS in Computer Science or related field

Tools

Terraform
AWS CloudFormation

Job description

Summary

You will join the PDO Site Reliability Engineering (Data Platform) team that looks after the reliability and operability of Guidewire’s data platform services, including large-scale data processing, analytics, and streaming capabilities that support our AI and Insight products. The team partners closely with product engineering, data platform, and security to design for reliability, build automation, and run services in production. As a Site Reliability Engineer – III (Data Platform), you will be a hands‑on contributor helping to run and evolve our big data stack on AWS (and potentially other public clouds/SaaS products), using software engineering to address infrastructure and application reliability challenges. You’ll collaborate with more senior SREs to improve incident response, harden critical data paths, and enable scalable, cost‑efficient operations, directly supporting PDO’s focus on operational excellence and AI/cloud/data platform adoption.

Job Description
What You Will Do
  • Collaborate with senior SREs to enhance incident response, reinforce critical data paths, and facilitate scalable, cost‑efficient operations in support of PDO’s objectives for operational excellence and AI/cloud/data platform adoption.
  • Support the maintenance and enhancement of production environment for data platform services to ensure high availability and performance for mission‑critical workloads.
  • Manage and optimize big data and streaming platforms like Kafka, Hadoop, Spark, and Hive on AWS, focusing on configuration, tuning, and daily operations.
  • Assist in defining SLOs, error budgets, capacity plans, and scaling strategies for data platform components to support new AI and data products.
  • Participate in on‑call rotations to maintain high availability of services, leveraging tools like PagerDuty to triage alerts and provide technical responses to incidents impacting data and analytics platforms.
  • Resolve production issues through cross‑team collaboration and adherence to established incident management practices.
  • Contribute to blameless post‑incident reviews to develop reliability improvements, runbooks, and automation.
  • Build and improve automation and tooling using Go, Python, or scripting to standardize deployment and troubleshooting for data services.
  • Support CI/CD pipelines using TeamCity or Github Actions to enable safe and frequent service deployments.
  • Utilize Infrastructure as Code, including Terraform and AWS CloudFormation, to maintain repeatable cloud infrastructure.
  • Operate Kubernetes‑based environments on AWS EKS, managing the lifecycle and scaling of containerized data services.
  • Implement progressive delivery strategies, such as blue/green and canary deployments, to minimize release risks.
What You Need To Succeed
Experience and Education
  • 4–6 years of relevant industry experience in Site Reliability Engineering, DevOps, Production Engineering, or similar roles supporting large‑scale distributed systems and/or data platforms.
  • BS/MS in Computer Science, Computer Engineering, Mathematics, or a related technical field, or equivalent practical experience.
Technical Skills
  • Experience deploying and operating services on AWS or Azure, including on‑call support.
  • Expertise with data and streaming platforms like Kafka, Hadoop, Spark, or Hive.
  • Proficiency in Go, Python, or Bash for automation; Java/Spring Boot knowledge is beneficial.
  • Skill in building tools and utilities using REST APIs or gRPC.
  • Proficiency with CI/CD tools such as TeamCity, GitHub Actions, or Jenkins.
  • Experience with Infrastructure as Code (Terraform or CloudFormation). Familiarity with Kubevela/Crossplane is a plus.
  • Working knowledge of Kubernetes (EKS) and Docker for deployment and resource management.
  • Familiarity with AWS services like RDS, EMR, Redshift, MSK, and ECS.
  • Experience with observability and logging tools like Datadog or ELK.
  • Understanding of distributed systems, networking, storage, and operating systems.
  • Familiarity with agile methodologies like Scrum and Kanban.
  • Ability to solve infrastructure problems using software engineering and participate in incident response.
  • Effective collaboration skills aligned with business priorities like AI and cloud adoption.
Bonus Points
  • Kubernetes/AWS certifications
  • Contributions to open source projects
Equal Opportunity & EEO Statement

Guidewire Software, Inc. is proud to be an equal opportunity and affirmative action employer. We are committed to an inclusive workplace, and believe that a diversity of perspectives, abilities, and cultures is a key to our success. Qualified applicants will receive consideration without regard to race, color, ancestry, religion, sex, national origin, citizenship, marital status, age, sexual orientation, gender identity, gender expression, veteran status, or disability. All offers are contingent upon passing a criminal history and other background checks where it’s applicable to the position.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (Application)
Senior Site Reliability Engineer (Application)

Guidewire Software • Kuala Lumpur

Hybrid
MYR 180,000 - 300,000
Software Engineer III
Software Engineer III

Guidewire Software • Kuala Lumpur

On-site
MYR 90,000 - 140,000
Software Engineer II
Software Engineer II

Guidewire Software • Kuala Lumpur

On-site
MYR 70,000 - 90,000
SRE III, Data Platform — AWS & Big Data
SRE III, Data Platform — AWS & Big Data

Guidewire Software • Kuala Lumpur

On-site
MYR 80,000 - 120,000
Senior Data Engineer
Senior Data Engineer

Involve Asia • Kuala Lumpur

On-site
MYR 70,000 - 90,000
Site Reliability Engineer (4024)
Site Reliability Engineer (4024)

Uncover • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Site Reliability Engineer
Site Reliability Engineer

Experian • Selangor

Hybrid
MYR 70,000 - 90,000
Hybrid working mode
Experian Care Allowance
Annual leaves
System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Malaysia • Kuala Lumpur

On-site
MYR 70,000 - 110,000
High-impact team environment
Opportunities for innovation
Influence engineering culture
Cloud Platform Support Engineer II (English /Japanese speaking)
Cloud Platform Support Engineer II (English /Japanese speaking)

Guidewire Software • Kuala Lumpur

Hybrid
MYR 89,000 - 156,000
Site Reliability Engineer (4024)
Site Reliability Engineer (4024)

GBG Plc • Kuala Lumpur

On-site
MYR 100,000 - 130,000