AI Infrastructure SRE — Kubernetes, GPU & Reliability

YTL AI Labs

Kuala Lumpur

On-site

MYR 120,000 - 200,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

YTL AI Labs in Malaysia seeks an AI Site Reliability Engineer to build, operate, and scale the core infra powering ILMU and the AI runtime layer that drives model serving, inference workloads, retrieval pipelines, and agent execution. You will be hands‑on across cloud, on‑prem GPU clusters, and hybrid deployments to ensure industry‑leading uptime.

This is a growth‑oriented role with learning opportunities in GPU workloads, AI infrastructure operations, and collaboration with AI research teams

Qualifications

  • 1–4 years in SRE, DevOps, infrastructure, systems administration, or equivalent roles.
  • Working knowledge of Linux systems and command‑line proficiency.
  • Hands‑on exposure to Kubernetes and containers (Docker) in production or lab settings.
  • Familiarity with at least one major cloud platform (AWS/Azure/GCP).
  • Basic scripting skills (Bash, Python, or similar).
  • Exposure to monitoring tools (Prometheus, Grafana, or similar).
  • Willingness to participate in on‑call rotations.

Responsibilities

  • Operate and maintain Kubernetes‑based environments across cloud and on‑prem GPU clusters under guidance from senior engineers.
  • Support deployment workflows and CI/CD pipelines, helping ensure safe, repeatable releases.
  • Maintain and improve operational runbooks, and contribute to automation that reduces manual toil.
  • Participate in on‑call rotations and assist in incident response and resolution.
  • Help build and maintain monitoring, logging, and alerting for model servers, vector DBs, agent frameworks, and platform.
  • APIs build and maintain dashboards that give teams real‑time visibility into system health.
  • Investigate alerts, triage issues, and upscale appropriately.
  • Assist in performance testing, benchmarking, and capacity tracking.

Skills

Linux
Kubernetes
Docker
Cloud platforms
Scripting (Bash/Python)
Monitoring (Prometheus, Grafana)
On-call experience

Tools

Kubernetes
Docker
Terraform
Prometheus
Grafana
GitHub Actions

Job description

YTL AI Labs in Malaysia seeks an AI Site Reliability Engineer to build, operate, and scale the core infra powering ILMU and the AI runtime layer that drives model serving, inference workloads, retrieval pipelines, and agent execution. You will be hands‑on across cloud, on‑prem GPU clusters, and hybrid deployments to ensure industry‑leading uptime.

This is a growth‑oriented role with learning opportunities in GPU workloads, AI infrastructure operations, and collaboration with AI research teams

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI SRE: Scalable, Reliable LLM Infra Architect
Senior AI SRE: Scalable, Reliable LLM Infra Architect

YTL AI Labs • Kuala Lumpur

On-site
MYR 180,000 - 360,000
AI Site Reliability Engineer
AI Site Reliability Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 120,000 - 200,000
Senior AI Site Reliability Engineer
Senior AI Site Reliability Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 180,000 - 360,000
AI Infra Engineer: GPU Cloud & ML Platform
AI Infra Engineer: GPU Cloud & ML Platform

Tencent • Kuala Lumpur

On-site
MYR 120,000 - 180,000
AI Infrastructure Engineer: GPU, Kubernetes & MLOps
AI Infrastructure Engineer: GPU, Kubernetes & MLOps

iSoftStone • Kuala Lumpur

On-site
MYR 60,000 - 100,000
Senior AI Platform Engineer: Build Scalable ML Infra
Senior AI Platform Engineer: Build Scalable ML Infra

Systems Limited • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Senior AI Infra Engineer – Production Reliability & Scale
Senior AI Infra Engineer – Production Reliability & Scale

Pertama Partners • Kuala Lumpur

On-site
MYR 80,000 - 120,000
Senior AI Engineer
Senior AI Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Production-Grade Full-Stack AI Engineer
Production-Grade Full-Stack AI Engineer

Resource Services Group X Pty Ltd • Kuala Lumpur

Hybrid
MYR 180,000 - 320,000
Hybrid working arrangements in Kuala L
Senior AI Product Engineer: Build National-Scale AI Systems
Senior AI Product Engineer: Build National-Scale AI Systems

YTL AI Labs • Kuala Lumpur

On-site
MYR 180,000 - 280,000