Senior Infrastructure Engineer, Data Compute Platform

Grab

Petaling Jaya

On-site

MYR 180,000 - 300,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Grab seeks a senior infrastructure engineer to own and scale the Data Compute Platform on Kubernetes (EKS) and AWS.

You will design provisioning, security, observability, and cost controls while driving reliable, cost-efficient infrastructure for data workloads. You will mentor engineers, review code, and contribute to SRE standards.

Qualifications

  • 3+ years of experience building and operating production infrastructure at scale.
  • Proficient in Go and/or Python; treat infrastructure as software.
  • Hands-on Kubernetes production experience with multi-tenancy.

Responsibilities

  • Design, build, and operate multi-tenant Kubernetes (EKS) platform.
  • Develop Kubernetes operators and custom resources to automate compute engines.
  • Lead infrastructure-as-code (Terraform) and CI/CD (GitLab) for AWS estate.
  • Design and implement identity, access, and security models across IAM and RBAC.
  • Build observability, alerting, capacity planning, and incident tooling.
  • Own compute cost efficiency and architectural improvements.

Skills

Go
Python
Kubernetes
AWS
Terraform
CI/CD
SRE basics
Code review

Education

Software Engineering, CS, or related undergraduate degree

Tools

Kubebuilder/Controller-runtime
GitLab CI
Jenkins
GitOps
Datadog
Prometheus
Grafana

Job description

Get to Know the Team

The Data Compute Platform team is an important contributor to Grab's data ecosystem, allowing growth through democratization of data at scale. We build and operate Grab's data infrastructure and efficient platform that supports internal data processes and company-wide data lake access. Our tech stack uses industry-leading distributed compute engines like Apache Spark, Ray, Trino, and Starrocks, orchestrated by Airflow and Michelangelo and backed by AWS S3. Our evolving Data Lake storage architecture uses modern open-source formats like Apache Iceberg and Delta in addition to traditional Apache Hive Parquet tables.

Underneath these engines sits a large, multi-tenant Kubernetes and AWS Infrastructure that we own end to end. This infrastructure includes the clusters, the operators, the autoscaling, the networking, the identity and access model, the observability, and the cost controls. These components work together to keep the platform fast, reliable, and affordable for thousands of pipelines and queries every day.

Get to Know the Role

You will be an important contributor to the infrastructure layer of the Data Compute Platform. This layer consists of the Kubernetes, AWS, and infrastructure-as-code foundations. Spark, Ray, Trino, Starrocks, Airflow, and Michelangelo run on these foundations. You will design how compute is provisioned, scaled, secured, observed and paid for, and you will drive the reliability and cost-efficiency of the platform as it grows. As a senior engineer, you will take ownership of well-scoped infrastructure projects from design through rollout and operation. You will also contribute to the team's engineering and SRE standards. Additionally, you will support other engineers through code review and knowledge sharing. You will also explore new developments in the cloud-native and data infrastructure space and integrate them into our ecosystem to the benefit of the data community at Grab.

You will report to our Data Engineering Manager II, and you will based onsite in our Petaling Jaya office.

The Critical Tasks You Will Perform

You will design, build, and operate the multi-tenant Kubernetes (EKS) platform. This platform runs Grab's workloads, including Spark, Ray, Trino, Starrocks, Airflow, and Michelangelo. Additionally, your responsibilities will include cluster lifecycle, node provisioning, and autoscaling, and scheduling and resource isolation.

You will build Kubernetes operators and custom resources (kubebuilder / controller-runtime, Go) that automate the provisioning and lifecycle of compute engines and their tenants.

You will lead the infrastructure-as-code (Terraform) and CI/CD (GitLab) that provision and change our AWS estate. This estate includes EKS, S3, IAM, RDS, VPC, and networking. You will drive it towards safe, reviewable, automated change.

You will design and implement the platform's identity, access and security model across AWS IAM, Kubernetes RBAC and service identities, working with the storage access and security teams.

You will build the observability, alerting, capacity planning, and incident tooling for the platform. You will contribute to the SRE practice, which includes SLOs, runbooks, on-call, and post-incident reviews. You will reduce toil and MTTR.

You will own compute cost efficiency: instance and storage strategy, spot and right-sizing, bin-packing, idle reclamation, and cost attribution back to tenants.

You will drive architectural improvements and migrations (for example engine version upgrades, cluster consolidation, new execution backends), managing the design, phased rollout and rollback plan with guidance from senior team members.

What Essential Skills You Will Need
  • Software Engineering, Computer Science, or related undergraduate degree.
  • You have 3 or more years of experience, with at least 2 years building and operating production infrastructure or platform services at scale.
  • You have programming proficiency in Go and/or Python, with the habit of treating infrastructure as software: tested, reviewed, versioned and automated.
  • You have deep, hands-on Kubernetes experience in production: cluster operations, scheduling, autoscaling, networking, storage, RBAC and multi-tenancy. We prefer experience building custom controllers or operators.
  • You have experience with AWS (EKS, EC2, S3, IAM, VPC) and infrastructure as code with Terraform.
  • You have proficiency in CI/CD tooling (GitLab CI, Jenkins or similar) and GitOps-style delivery.
  • You have solid SRE fundamentals: observability (Datadog, Prometheus, Grafana or equivalent), SLOs, incident management and capacity planning, with experience running reliable services.
Skills that are Good to have
  • Working knowledge of at least one of Spark, Ray, Airflow, Trino or Starrocks, and a appetite to learn how distributed data engines behave on Kubernetes.
  • Experience running Apache Spark on Kubernetes at scale (Spark Operator, dynamic allocation, shuffle services) and tuning its interaction with the
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Infrastructure Engineer, Data Compute Platform
Senior Infrastructure Engineer, Data Compute Platform

PLT engineering GmbH • Selangor

On-site
MYR 150,000 - 230,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

PLT Engineering • Selangor

On-site
MYR 180,000 - 320,000
Senior Data Compute Platform Infra Engineer
Senior Data Compute Platform Infra Engineer

PLT engineering GmbH • Selangor

On-site
MYR 150,000 - 230,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

GrabTaxi Holdings Pte. Ltd. • Petaling Jaya

On-site
MYR 180,000 - 280,000
Term Life Insurance
Medical Insurance
GrabFlex benefits
+5
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Grab • Petaling Jaya

On-site
MYR 180,000 - 280,000
Term Life Insurance
Medical Insurance
GrabFlex benefits package
+5
Senior Software Engineer - Deployment Platform
Senior Software Engineer - Deployment Platform

Grab • Petaling Jaya

On-site
MYR 180,000 - 280,000
Term Life Insurance
Medical Insurance
GrabFlex benefits
+5
Lead Software Engineer, Managed Kubernetes Platform
Lead Software Engineer, Managed Kubernetes Platform

Grab • Petaling Jaya

On-site
MYR 100,000 - 150,000
Term Life Insurance
Comprehensive Medical Insurance
GrabFlex benefits package
+4
Lead Software Engineer, Managed Kubernetes Platform
Lead Software Engineer, Managed Kubernetes Platform

GrabTaxi Holdings Pte. Ltd. • Petaling Jaya

On-site
MYR 80,000 - 110,000
Term Life Insurance
Comprehensive Medical Insurance
Flexible work arrangements
+3
Senior Software Engineer - Deployment Platform
Senior Software Engineer - Deployment Platform

PLT Engineering • Selangor

On-site
MYR 150,000 - 240,000
Senior Data Engineer
Senior Data Engineer

Involve Asia • Kuala Lumpur

On-site
MYR 70,000 - 90,000