AI Infrastructure & Orchestration Lead

SNS Network (M) Sdn. Bhd.

Petaling Jaya

On-site

MYR 180,000 - 260,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job description

AI Engineer

We are seeking a highly skilled and versatile AI Infrastructure & Systems Specialist to drive the deployment, orchestration, and management of our AI computing infrastructure. This role is critical in ensuring the smooth operation of ourGPUaaS platform, optimizing AI workloads, and supporting customers in their AI adoption journey. The ideal candidate should have expertise in GPU systems, AI orchestration platforms, software engineering for AI, and technical leadership.

Key Responsibilities

  1. GPU Systems Specialist
  • Deploy, configure, and manageNVIDIA H100-poweredGPU servers and networking infrastructure (InfiniBand).
  • Optimize GPU performance for AI/ML workloads and troubleshoot hardware/software issues.
  • Collaborate with NVIDIA and Dell teams for system optimization and technical support.
  • Validate the GPU fractions, perform necessary checks and upgrades on GPU device drivers and Nvidia libraries.
  1. Orchestration & Virtualization Specialist
  • Implement and manageRUN.AI, Kubernetes, and NVIDIA MIGto optimize GPU resource allocation.
  • Ensure seamlessmulti-tenant AI workload managementand scaling strategies.
  • Automate AI/ML pipeline orchestration for efficient resource utilization.
  • Monitor the Services exposed via Kubernetes orchestration, enable GitOps model for the AI/ML workflows and ensure the K8s cluster works optimally.
  • Monitor the Cluster usage, GPU quotas, Storage utilization and build performance reports on regular basis to assess the functionality of the K8s GPU stack.
  1. AI Software Engineer
  • Work with AI teams to enable model training and fine-tuning usingPyTorch, TensorFlow, and RAPIDS.
  • Develop and optimize AI/ML workflows on high-performance computing (HPC) environments.
  • IntegrateAI frameworks with cloud and on-prem GPU clusters.
  1. Technical Manager
  • Serve as theprimary technical advisorto customers, ensuring seamless AI deployment.
  • Conduct technical workshops, bootcamps, and onboarding sessions for users.
  • Collaborate with universities, enterprises, and startups to drive AI adoption.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer SNS Network Right Choice with the Right People
AI Engineer SNS Network Right Choice with the Right People

SNS Network (M) Sdn. Bhd. • Petaling Jaya

On-site
MYR 180,000 - 260,000
Technical Manager - GPU Cloud & AI Infrastructure
Technical Manager - GPU Cloud & AI Infrastructure

Risewave Consulting, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 280,000
AI DevOps Engineer (MLOps & Cloud)
AI DevOps Engineer (MLOps & Cloud)

DXC Technology Inc. • Petaling Jaya

On-site
MYR 100,000 - 150,000
IEGG - AI Infra Engineer (3rd Party Contract -1 Year Renewable)
IEGG - AI Infra Engineer (3rd Party Contract -1 Year Renewable)

Tencent • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Senior Engineer (Infra)
Senior Engineer (Infra)

Pertama Partners • Kuala Lumpur

On-site
MYR 80,000 - 120,000
Senior Linux Infra Engineer: Bare Metal, GPU & AI
Senior Linux Infra Engineer: Bare Metal, GPU & AI

Uvation • Malaysia

Remote
MYR 180,000 - 240,000
Lead, Forward Deployed Engineer
Lead, Forward Deployed Engineer

SARAWAK ARTIFICIAL INTELLIGENCE CENTRE SDN BHD • Kuching

On-site
MYR 150,000 - 210,000
Principal Cloud Architect – HPC/GPU & AI Platform Solutions
Principal Cloud Architect – HPC/GPU & AI Platform Solutions

Ll Oefentherapie • Kuala Lumpur

On-site
MYR 240,000 - 420,000
Forward Deployed Architect
Forward Deployed Architect

SARAWAK ARTIFICIAL INTELLIGENCE CENTRE SDN BHD • Kuching

On-site
MYR 250,000 - 500,000
AI/ML Engineer III
AI/ML Engineer III

Astreya • Kuala Lumpur

On-site
MYR 120,000 - 240,000