Artificial Intelligence Engineer

iSoftStone

Kuala Lumpur

On-site

MYR 60,000 - 100,000

Full time

18 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

iSoftStone is seeking an AI Infrastructure Engineer in Kuala Lumpur to build and operate large-scale GPU AI infra, including AWS cloud environments and GPU Kubernetes clusters. You will optimize resource usage, support ML workloads, and drive automation.

The role requires Python programming, Linux, Docker, Kubernetes, and cloud experience (preferably AWS). English and Chinese language skills are essential for global collaboration and project success.

Qualifications

  • Bachelor’s degree as listed in requirements.
  • Minimum 1 year of relevant AI infrastructure related experience.
  • Proficiency in English and Chinese for global collaboration.
  • Strong Python or backend development language experience.
  • Practical experience with Linux, Docker, and Kubernetes (K8s).
  • Hands-on experience with cloud platforms, preferably AWS.
  • Solid experience with AI model training, inference deployment, or ML production environments.
  • Experience with PyTorch, DeepSpeed, Megatron, and vLLM is preferred.
  • Understanding of LLM training/inference workflows and distributed training concepts.
  • Experience with GPU computing environments, resource scheduling, or AI infrastructure operations is a plus.

Responsibilities

  • Operate and maintain GPU-based AI infrastructure including AWS and bare-metal GPU Kubernetes clusters.
  • Support GPU cluster deployment, configuration, monitoring, capacity management, troubleshooting, and incident resolution.
  • Optimize GPU resource utilization and workload scheduling for efficiency.
  • Develop tools for AI computing resource management, scheduling, monitoring, and automation.
  • Support ML teams in deploying AI model training and inference workloads.
  • Troubleshoot end-to-end ML workflows from data preparation to model serving.
  • Collaborate across teams to ensure scalable AI infrastructure.

Skills

Python programming
Linux
Docker
Kubernetes
Cloud platforms (AWS)
English and Chinese
AI/ML workflows

Education

Bachelor’s degree in Computer Science / Software Engineering / Artificial Intelligence

Tools

Docker
Kubernetes (K8s)
AWS

Job description

We are a leading global technology group with a strong presence across cloud computing, AI, mobile gaming, social media, and enterprise technology. Our platforms serve millions of users and businesses worldwide, with a strong focus on innovation, scalability, reliability, and security.

As we continue to expand our AI capabilities, we are looking for an AI Infrastructure Engineer to help build and operate the infrastructure powering AI model training and inference at scale.

GPU & AI Infrastructure
  • Operate and maintain GPU-based AI infrastructure, including AWS cloud environments and bare-metal GPU Kubernetes (K8s) clusters.
  • Support GPU cluster deployment, configuration, monitoring, capacity management, troubleshooting, and incident resolution.
  • Optimize GPU resource utilization, workload scheduling, and infrastructure efficiency.
  • Develop and maintain tools for AI computing resource management, scheduling, monitoring, and automation.
  • Support Machine Learning / Algorithm teams in deploying and running AI model training and inference workloads.
  • Understand and troubleshoot end-to-end ML workflows, from data/model preparation and training to model serving and inference.
  • Support distributed and large-scale AI workloads running across GPU clusters.
  • Identify infrastructure bottlenecks and optimize AI workloads for performance, reliability, and resource efficiency.
AI Framework & ML Platform
  • Work with AI/ML frameworks and tools such as PyTorch, DeepSpeed, Megatron, and vLLM.
  • Build and maintain infrastructure and workflows supporting AI model training and inference.
  • Support LLM training and inference workflows, including distributed GPU workloads.
  • Collaborate with algorithm teams to improve the reliability and scalability of AI development and production environments.
  • Deploy and manage AI workloads on Kubernetes (K8s) environments.
  • Build, maintain, and optimize CI/CD pipelines for AI/ML applications and infrastructure.
  • Automate deployment, testing, monitoring, and operational workflows.
  • Improve the reliability and efficiency of AI services through infrastructure automation and DevOps/MLOps practices.
Requirements:
  • Bachelor’s degree in Computer Science, Software Engineering, Artificial Intelligence, or related fields.
  • Minimum 1 year of relevant experience in AI Infrastructure, MLOps, ML Engineering, Backend Engineering, Cloud Engineering, DevOps, SRE, or a related field.
  • Proficiency in both English and Chinese languages, as this role is required for collaboration with global team.
  • Strong programming experience in Python or other backend development languages.
  • Practical experience with Linux systems, Docker, and Kubernetes (K8s).
  • Must have hands-on experience with cloud platforms, preferably AWS.
  • Solid project experience with AI model training, inference deployment, or ML production environments is highly preferred.
  • Practical experience with AI/ML frameworks such as: PyTorch, DeepSpeed, Megatron, vLLM
  • Strong Understanding of Large Language Model (LLM) training/inference workflows and distributed training concepts
  • With project experience with GPU computing environments, resource scheduling, or AI infrastructure operation is a strong plus point.
Preferred Qualifications
  • Working experience in MLOps or ML Platform Engineering.
  • Experience supporting GPU clusters or AI computing platforms.
  • Strong understanding of distributed training strategies, including data parallelism, tensor parallelism, and pipeline parallelism.
  • Strong problem-solving skills and ability to troubleshoot production issues.
  • Passionate about AI technologies and capable of leveraging AI tools to improve engineering efficiency.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

IEGG - AI Infra Engineer (3rd Party Contract -1 Year Renewable)
IEGG - AI Infra Engineer (3rd Party Contract -1 Year Renewable)

Tencent • Kuala Lumpur

On-site
MYR 120,000 - 180,000
AI DevOps Engineer (MLOps & Cloud)
AI DevOps Engineer (MLOps & Cloud)

DXC Technology Inc. • Petaling Jaya

On-site
MYR 100,000 - 150,000
AI Infrastructure & Orchestration Lead
AI Infrastructure & Orchestration Lead

SNS Network (M) Sdn. Bhd. • Petaling Jaya

On-site
MYR 180,000 - 260,000
AI Engineer SNS Network Right Choice with the Right People
AI Engineer SNS Network Right Choice with the Right People

SNS Network (M) Sdn. Bhd. • Petaling Jaya

On-site
MYR 180,000 - 260,000
AI Engineer
AI Engineer

Lenovo • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Lead AI/ML Cloud Engineer_2026
Lead AI/ML Cloud Engineer_2026

Private Advertiser • Selangor

On-site
MYR 120,000 - 180,000
Latest technologies
Cloud-based AI solution
International team
+1
Technical Manager - GPU Cloud & AI Infrastructure
Technical Manager - GPU Cloud & AI Infrastructure

Risewave Consulting, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Lead, Forward Deployed Engineer
Lead, Forward Deployed Engineer

SARAWAK ARTIFICIAL INTELLIGENCE CENTRE SDN BHD • Kuching

On-site
MYR 150,000 - 210,000
AI Infra Engineer: GPU Cloud & ML Platform
AI Infra Engineer: GPU Cloud & ML Platform

Tencent • Kuala Lumpur

On-site
MYR 120,000 - 180,000
AI Infrastructure Engineer: GPU, Kubernetes & MLOps
AI Infrastructure Engineer: GPU, Kubernetes & MLOps

iSoftStone • Kuala Lumpur

On-site
MYR 60,000 - 100,000