Senior Site Reliability Engineer (Data Platform)

Guidewire Software

Kuala Lumpur

On-site

MYR 240,000 - 360,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Guidewire invites a Senior Site Reliability Engineer (Data Platform) to lead reliability across our data stack on AWS, building automation, scalable cloud infrastructure, and robust observability. You will drive CI/CD automation, blue/green deployments, and incident response improvements while collaborating with product, data, and security teams to support AI and analytics workloads.

The role emphasizes governance, cost-aware design, and continuous improvement, with a focus on reliability and

Qualifications

  • 8+ years of experience in Site Reliability Engineering, DevOps, or similar roles.
  • Formal degree in CS/Engineering or equivalent practical experience.

Responsibilities

  • Design and implement self-service automation and tooling to streamline deployment, operations, and troubleshooting for data platform services.
  • Implement and improve CI/CD pipelines to support safe, frequent deployments with automated checks.
  • Use Infrastructure as Code to build, harden, and maintain repeatable cloud infrastructure for data workloads.
  • Operate and improve Kubernetes-based environments (AWS EKS) for containerized data services.
  • Apply blue/green and canary deployments and support chaos engineering for resilience validation.
  • Collaborate on capacity planning and cost-aware cloud design across compute, storage, and networking.
  • Build end-to-end observability using Datadog and ELK, with metrics, traces, and dashboards.
  • Develop dashboards and alerts for data pipeline health and platform performance.
  • Analyze operational data to identify reliability risks and inform roadmaps.
  • Partner with product engineering, data platform, security, and other SRE teams on architectural improvements.

Skills

Automation tooling
CI/CD pipelines
Infrastructure as Code
Kubernetes & Docker
Cloud platforms (AWS)
Monitoring/Observability

Education

BS/MS in CS/Engineering or equivalent

Tools

Terraform
AWS CloudFormation
Datadog
ELK
TeamCity
Github Actions
Kafka
Hadoop
Spark

Job description

Summary
The Team and the Opportunity

You will join the PDO Site Reliability Engineering (Data Platform) team that owns the reliability and operability of Guidewire’s data platform services, including large-scale data processing, analytics, and streaming capabilities that underpin our AI and Insight products.

This team partners closely with product engineering, data platform, and security to design for reliability, build automation, and run services in production.

As a Senior Site Reliability Engineer (Data Platform), you will be a technical leader in running and evolving our big data stack on AWS (and potentially other public clouds), using software engineering to solve infrastructure and application reliability challenges. You’ll help advance
PDO’s priorities by improving incident response, hardening critical data paths, and enabling scalable, cost-efficient operations for our customers’ most important workloads.

Job Description
What You Will Do
  • Design and implement self-service automation and tooling (in Go, Python, or scripting languages) to standardize and streamline deployment, operations, and troubleshooting for data platform services.

  • Implement and improve CI/CD pipelines (e.g., TeamCity, Github Actions) to support safe, frequent deployments, including gate promotion and automated quality checks.

  • Use Infrastructure as Code (e.g., Terraform, AWS CloudFormation) to build, harden, and maintain repeatable cloud infrastructure for data and analytics workloads.

  • Operate and improve Kubernetes-based environments (AWS EKS), including deployment, scaling, and lifecycle management of containerized data services (e.g., Docker-packaged microservices, streaming jobs).

  • Apply progressive delivery strategies such as blue/green and canary deployments, and support chaos engineering experiments to validate resilience and recovery mechanisms.

  • Collaborate on capacity planning and cost-aware design for cloud resources across compute, storage, and networking layers for data-intensive systems.

  • Build and refine end-to-end observability for the data platform using monitoring and logging tools (e.g., Datadog, ELK), including metrics, traces, and logs.

  • Develop meaningful dashboards and alerts to provide clear visibility into data pipeline health, customer experience, and platform performance.

  • Analyze operational data to identify reliability risks and bottlenecks, feeding insights into the roadmap and reliability backlogs.

  • Partner with product engineering, data platform, security, and other SRE teams to define and implement improvements in service architecture and operational practices that support PDO’s AI, cloud, and data platform priorities.

  • Advocate for reliability, resilience, and operational excellence in design reviews, readiness assessments, and release planning.

  • Contribute to a positive, inclusive work environment based on accountability, continuous learning, and psychological safety, consistent with Guidewire’s culture of determination, collaboration, continuous improvement, and bravery.

What You Need to Succeed
Experience and Education
  • 8+ years of relevant industry experience in Site Reliability Engineering, DevOps, Production Engineering, or similar roles supporting large-scale distributed systems and data platforms.

  • BS/MS in Computer Science, Computer Engineering, Mathematics, or equivalent practical experience.

Technical Skills
  • Strong experience with continuous deployment and operation of cloud services on public cloud (AWS), including production support and on-call.

  • Hands-on experience running data platforms using big data and streaming technologies such as Kafka, Hadoop, Spark, and Hive on the public cloud.

  • Proficiency in at least one of Java, Go, or Python, and solid skills with scripting languages to build tools, automation, and integrations.

  • Experience building and operating microservices, including REST APIs and/or gRPC services.

  • Solid experience with CI/CD tools (e.g., TeamCity, Github Actions) for automated builds, tests, and deployments, including promotion gates.

  • Strong experience with Infrastructure as Code tools such as Terraform and AWS CloudFormation for provisioning and managing cloud infrastructure. Familiarity with Kubevela/Crossplane is a plus

  • Practical knowledge of Kubernetes (e.g., AWS EKS) and Docker, including deployment patterns, service discovery, and resource management.

  • Familiarity with AWS services relevant to data and distributed systems, such as RDS, EMR, Redshift, MSK (Managed Streaming for Kafka), ECS, SNS, and SQS. Expertise with monitoring, logging, and observability tools (e.g., Datadog, ELK) to instrument services and build actionable alerts and dashboards.

  • Deep understanding of distributed systems fundamentals, networking, storage, operating systems, and how they interact in complex multi-tier environments.

  • Knowledge of capacity planning, scalability, and resilience patterns (including blue/green and canary deployments, and chaos engineering concepts).

Operational and Problem-Solving Skills
  • Demonstrated experience solving infrastructure and application problems using software engineering approaches rather than only manual operations.

  • Familiarity with agile methodologies like Scrum and Kanban.

  • Strong analytical and troubleshooting skills for complex, distributed, multi-service environments.

  • Experience with on-call, incident response (e.g., PagerDuty), and post-incident review processes, with a bias for learning and continuous improvement.

Ways of Working
  • Ability to collaborate effectively with other engineering, data, and operations teams to understand their systems and help improve them.

  • A big-picture perspective on systems, tools, and customer value, aligning technical decisions with PDO’s priorities around operational excellence, AI, cloud, and data platform adoption.

  • Comfort with agile development methodologies and iterative delivery in a highly collaborative environment.

  • Eagerness to learn, experiment, and grow—staying current with emerging technologies across cloud, data, and SRE practices, and applying them thoughtfully where they add real value.

Bonus Points
  • Kubernetes/AWS certifications

  • Contributions to open source projects

#LI-AA1

About Guidewire

Guidewire is the platform P&C insurers trust to engage, innovate, and grow efficiently. We combine digital, core, analytics, and AI to deliver our platform as a cloud service. More than 540+ insurers in 40 countries, from new ventures to the largest and most complex in the world, run on Guidewire.

As a partner to our customers, we continually evolve to enable their success. We are proud of our unparalleled implementation track record with 1600+ successful projects, supported by the largest R&D team and partner ecosystem in the industry. Our Marketplace provides hundreds of applications that accelerate integration, localization, and innovation.

For more information, please visit www.guidewire.com and follow us on Twitter: @Guidewire_PandC.

Guidewire Software, Inc. is proud to be an equal opportunity and affirmative action employer. We are committed to an inclusive workplace, and believe that a diversity of perspectives, abilities, and cultures is a key to our success. Qualified applicants will receive consideration without regard to race, color, ancestry, religion, sex, national origin, citizenship, marital status, age, sexual orientation, gender identity, gender expression, veteran status, or disability. All offers are contingent upon passing a criminal history and other background checks where it’s applicable to the position.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Global Delivery Associate Consultant, Guidewire Cloud
Global Delivery Associate Consultant, Guidewire Cloud

Guidewire Software • Kuala Lumpur

On-site
Cloud Platform Support Engineer II (English /Japanese speaking)
Cloud Platform Support Engineer II (English /Japanese speaking)

Guidewire Software • Kuala Lumpur

Hybrid
MYR 89,000 - 156,000
Senior Data Platform SRE — Cloud, Kubernetes & Analytics
Senior Data Platform SRE — Cloud, Kubernetes & Analytics

Guidewire Software • Kuala Lumpur

On-site
MYR 240,000 - 360,000
Associate Consultant, Global Delivery
Associate Consultant, Global Delivery

Guidewire Software • Kuala Lumpur

On-site
Site Reliability Engineer
Site Reliability Engineer

Experian Group • Cyberjaya

Hybrid
MYR 180,000 - 240,000
Great compensation
Discretionary bonus
Hybrid work arrangement
+1
Consultant - Policy/Billing Center
Consultant - Policy/Billing Center

Guidewire Software • Kuala Lumpur

On-site
MYR 120,000 - 210,000
Senior Director, Platform Engineering - SDLC & Cloud Native Foundations
Senior Director, Platform Engineering - SDLC & Cloud Native Foundations

Experian • Selangor

Hybrid
MYR 726,000 - 968,000
Great compensation package
Hybrid work arrangement
Equal opportunities employer
Site Reliability Engineer
Site Reliability Engineer

Experian Asia Pacific • Cyberjaya

Hybrid
MYR 180,000 - 240,000
Great compensation package
Hybrid work arrangement
Equal opportunities employer
Azure Infrastructure & Data Engineer
Azure Infrastructure & Data Engineer

S&P Global, Inc. • Malaysia

On-site
MYR 120,000 - 240,000
Health care coverage
Flexible time off
Continuous learning opportunities
+3
Senior Data Engineer
Senior Data Engineer

Involve Asia • Kuala Lumpur

On-site
MYR 70,000 - 90,000