AI Infra Engineer (ML Platform, AI Native production, Algorithm, cutting-edge technology, multinational company)

DADACONSULTANTS PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Our client, a well-funded AI company, designs and runs large-scale compute infrastructure powering frontier model training and inference. They seek an infrastructure engineer to design, operate, and improve their GPU cluster platform, a deeply technical role at the junction of distributed systems and ML platform engineering.

You will own the compute platform, optimize GPU resource sharing, tackle bottlenecks across compute, storage and networking, and build automation and observability.

Qualifications

  • 3+ years in large-scale infrastructure or distributed platforms.
  • Strong knowledge of Kubernetes and containerized environments.
  • Solid Linux systems, resource scheduling, storage, networking.
  • Exposure to GPU clusters, distributed training is a plus.

Responsibilities

  • Own the compute platform that supports large-scale AI model training and serving.
  • Improve GPU resource sharing and scheduling across teams and workloads.
  • Solve infrastructure bottlenecks across compute, storage, networking, and distributed communication.
  • Develop platform capabilities that make cluster operations more automated, observable, and reliable.
  • Work with AI engineering teams to improve workload performance and overall infrastructure efficiency.

Skills

Kubernetes
Linux
Distributed systems
Go
Python
C++
HPC networking

Job description

Our client is a well-funded AI company operating at significant scale, building and running the large-scale compute infrastructure that powers frontier model training and inference. They are looking for a strong infrastructure engineer to design, operate, and continuously improve their GPU cluster platform. This is a high-impact, deeply technical role sitting at the intersection of distributed systems, ML platform engineering, and large-scale operations — ideal for engineers who want to work close to the metal on some of the most demanding infrastructure challenges in the industry.

What You’ll Do
  • Own the compute platform that supports large-scale AI model training and serving.
  • Improve how GPU resources are shared and scheduled across teams and workloads.
  • Solve infrastructure bottlenecks across compute, storage, networking, and distributed communication.
  • Develop platform capabilities that make cluster operations more automated, observable, and reliable.
  • Work with AI engineering teams to improve workload performance and overall infrastructure efficiency.
What We’re Looking For
  • 3+ years of experience building large-scale infrastructure or distributed platforms.
  • Strong knowledge of Kubernetes and containerized environments, with practical experience operating production clusters.
  • Solid understanding of Linux systems, resource scheduling, storage, and networking.
  • Exposure to GPU clusters, distributed training, NCCL/RDMA or HPC networking is highly preferred.
  • Strong programming ability in Go, Python, or C++.
  • Experience supporting large-scale ML/AI workloads is a strong advantage.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer — GPU Cluster & ML Platform
Senior AI Infra Engineer — GPU Cluster & ML Platform

DADACONSULTANTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SWAPETECH PTE. LTD. • Singapore

On-site
SGD 120,000 - 170,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Infrastructure Engineer
AI Infrastructure Engineer

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
Software Engineer, ML Dev Enablement
Software Engineer, ML Dev Enablement

MOTIONAL SINGAPORE PTE. LIMITED • Singapore

On-site
SGD 80,000 - 130,000
AI Infra Engineer: GPU HPC Clusters & Orchestration
AI Infra Engineer: GPU HPC Clusters & Orchestration

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Architect: GPU Clusters & HPC
AI Infra Architect: GPU Clusters & HPC

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Systems Infra Engineer - Multi-GPU HPC
AI Systems Infra Engineer - Multi-GPU HPC

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000