AI Site Reliability Engineer

YTL AI Labs

Kuala Lumpur

On-site

MYR 120,000 - 200,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

YTL AI Labs in Malaysia seeks an AI Site Reliability Engineer to build, operate, and scale the core infra powering ILMU and the AI runtime layer that drives model serving, inference workloads, retrieval pipelines, and agent execution. You will be hands‑on across cloud, on‑prem GPU clusters, and hybrid deployments to ensure industry‑leading uptime.

This is a growth‑oriented role with learning opportunities in GPU workloads, AI infrastructure operations, and collaboration with AI research teams

Qualifications

  • 1–4 years in SRE, DevOps, infrastructure, systems administration, or equivalent roles.
  • Working knowledge of Linux systems and command‑line proficiency.
  • Hands‑on exposure to Kubernetes and containers (Docker) in production or lab settings.
  • Familiarity with at least one major cloud platform (AWS/Azure/GCP).
  • Basic scripting skills (Bash, Python, or similar).
  • Exposure to monitoring tools (Prometheus, Grafana, or similar).
  • Willingness to participate in on‑call rotations.

Responsibilities

  • Operate and maintain Kubernetes‑based environments across cloud and on‑prem GPU clusters under guidance from senior engineers.
  • Support deployment workflows and CI/CD pipelines, helping ensure safe, repeatable releases.
  • Maintain and improve operational runbooks, and contribute to automation that reduces manual toil.
  • Participate in on‑call rotations and assist in incident response and resolution.
  • Help build and maintain monitoring, logging, and alerting for model servers, vector DBs, agent frameworks, and platform.
  • APIs build and maintain dashboards that give teams real‑time visibility into system health.
  • Investigate alerts, triage issues, and upscale appropriately.
  • Assist in performance testing, benchmarking, and capacity tracking.

Skills

Linux
Kubernetes
Docker
Cloud platforms
Scripting (Bash/Python)
Monitoring (Prometheus, Grafana)
On-call experience

Tools

Kubernetes
Docker
Terraform
Prometheus
Grafana
GitHub Actions

Job description

At YTL AI Labs, we build sovereign AI models that perform on par with the world’s best- while staying grounded in local needs, values, and context. Our flagship model, ILMU, is designed to be culturally aware, contextually intelligent, and fluent in Bahasa Melayu, delivering cutting-edge solutions that empower Malaysian businesses with intelligence that truly understands the market and the people they serve.As pioneers of sovereign AI, we believe every nation should have the power to shape its own intelligenc - guided by its people, priorities, and principles.

About the Role

As an AI Site Reliability Engineer, you will build, operate, and scale the core infrastructure powering ILMU and the AI runtime layer that drives model serving, inference workloads, retrieval pipelines, and agent execution. You will be hands‑on in ensuring the reliability and performance of our infrastructure across the cloud, on‑prem GPU clusters, and hybrid deployments, so that our LLM inference, agentic workflows, and platform services run with industry‑leading uptime and efficiency.

This is a hands‑on role with significant room to grow. You will learn how production‑grade LLM infrastructure is built and operated, from Kubernetes and observability pipelines to GPU clusters and model‑serving platforms, while taking real ownership of monitoring, automation, and incident response tasks from day one.

Key Responsibilities:

Infrastructure Operations

  • Operate and maintain Kubernetes‑based environments across cloud and on‑prem GPU clusters under guidance from senior engineers
  • Support deployment workflows and CI/CD pipelines, helping ensure safe, repeatable releases
  • Maintain and improve operational runbooks, and contribute to automation that reduces manual toil
  • Participate in on‑call rotations and assist in incident response and resolution

Observability & Monitoring

  • Help build and maintain monitoring, logging, and alerting for model servers, vector DBs, agent frameworks, and platform
  • APIsBuild and maintain dashboards that give teams real‑time visibility into system health
  • Investigate alerts, triage issues, and upscale appropriately
  • Assist in performance testing, benchmarking, and capacity tracking

Reliability & Continuous Improvement

  • Help track SLIs/SLOs and flag services at risk of breaching targets
  • Contribute to postmortems and follow through on action items in a blameless culture
  • Identify recurring operational issues and propose fixes or automation
  • Follow security and access‑control best practices across all environments
  • Work closely with senior SREs, platform engineering, and AI research teams
  • Learn AI infrastructure operations: GPU workloads, inference serving, and retrieval pipelines

Skills & Qualifications

Must‑Have

  • 1–4 years in SRE, DevOps, infrastructure, systems administration, or equivalent roles (fresh graduates with strong relevant projects or internships considered for junior level)
  • Working knowledge of Linux systems and command‑line proficiency
  • Hands‑on exposure to Kubernetes and containers (Docker), in production or substantial lab/project settings
  • Familiarity with at least one major cloud platform (AWS/Azure/GCP)
  • Basic scripting skills (Bash, Python, or similar)
  • Exposure to monitoring tools (Prometheus, Grafana, or similar)
  • A strong learning mindset and willingness to participate in on‑call roations

Bonus

  • Exposure to on‑prem infrastructure (SAN, Proxmox, firewalls, switches, routers)
  • Exposure to GitOps/DevOps tooling (ArgoCD, Terraform, LGTM monitoring stacks)
  • Interest in or exposure to LLM inference, model serving (vLLM, SGLang, Triton, TGI), or benchmark testing
  • Familiarity with CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins)
  • Understanding of basic networking (VPCs, load balancers, DNS)

What Success Looks Like

  • You independently handle routine operational tasks, alerts, and first‑line incident triage within your first few months
  • Dashboards and runbooks you maintain are accurate, useful, and trusted by the team
  • You contribute meaningful automation that reduces repetitive manual work
  • You grow steadily in AI infrastructure expertise - GPU operations, inference serving, observability, with a clear path toward senior SRE responsibilities
  • You participate constructively in postmortems and help close out reliability action items
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Site Reliability Engineer
Senior AI Site Reliability Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 180,000 - 360,000
AI Infrastructure SRE — Kubernetes, GPU & Reliability
AI Infrastructure SRE — Kubernetes, GPU & Reliability

YTL AI Labs • Kuala Lumpur

On-site
MYR 120,000 - 200,000
Senior AI SRE: Scalable, Reliable LLM Infra Architect
Senior AI SRE: Scalable, Reliable LLM Infra Architect

YTL AI Labs • Kuala Lumpur

On-site
MYR 180,000 - 360,000
Senior AI Engineer
Senior AI Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Model Performance Engineer
Model Performance Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Full Stack AI Engineer
Full Stack AI Engineer

RSGx • Kuala Lumpur

Hybrid
MYR 180,000 - 240,000
Hybrid work
KL office nearby
AI Product Engineer
AI Product Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Full Stack AI Engineer
Full Stack AI Engineer

Resource Services Group X Pty Ltd • Kuala Lumpur

Hybrid
MYR 180,000 - 320,000
Hybrid working arrangements in Kuala L
AI Solutions Engineer
AI Solutions Engineer

techstreet • Petaling Jaya

On-site
MYR 60,000 - 110,000
AI Systems Support & Quality Assurance Engineer
AI Systems Support & Quality Assurance Engineer

Jobstreet Malaysia • Selangor

On-site
MYR 30,000 - 42,000