AI Engineer (ML Systems & Infrastructure)

SWAPETECH PTE. LTD.

Singapore

On-site

SGD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SWAPETECH PTE. LTD. seeks exceptional AI Engineers to build the next generation of AI infrastructure and MLSys systems. You will focus on large-scale distributed systems, Kubernetes-based AI infrastructure, RDMA networking, KV Cache, GPUs, and CUDA kernel optimisation.

You will collaborate with AI researchers, infra architects, and platform teams to maximise efficiency, scalability, and reliability of AI workloads.

Qualifications

  • Bachelor's degree or above in CS/SE/EE or related field.
  • Strong software engineering and system design capability.
  • Experience with distributed systems and GPU infrastructure.

Responsibilities

  • Design, deploy, and operate large-scale Kubernetes-based AI infrastructure.
  • Develop cluster governance, scheduling, resource isolation, and multi-tenancy.
  • Build and optimize GPU orchestration platforms and high-speed networking.
  • Improve cluster utilization, reliability, and operational efficiency.
  • Develop observability platforms, dashboards, and SLO-driven ops.

Skills

Strong software engineering
System design
C++/Go/Python/Rust
GPU performance engineering
Distributed systems
Networking
Kubernetes
CUDA
NCCL
Observability

Education

Bachelor's degree or above in Computer Science, Software Engineering, Electrical Engineering, or related fields

Tools

Kubernetes
CUDA
NCCL
TensorRT
Triton
Ray
Volcano
Kueue

Job description

About the Role

We are looking for exceptional AI Engineers to build the next generation of AI infrastructure and Machine Learning Systems(MLSys).

This role focuses on large-scale system infrastructure rather than model research. You will work on the core foundations that power large-scale AI training and inference systems, including Kubernetes cluster management, RDMA networking, unified KV Cache architecture, observability platforms, distributed systems, GPU orchestration, and CUDA kernel optimisation.

You will collaborate closely with AI researchers, infrastructure architects, networking engineers, and platform teams to maximize the efficiency, scalability, and reliability of AI systems.

Key Responsibilities
AI Infrastructure & Kubernetes
  • Design, deploy, and operate large-scale Kubernetes-based AI infrastructure.
  • Develop cluster governance frameworks, scheduling policies, resource isolation, and multi-tenancy capabilities.
  • Build and optimize GPU orchestration platforms using Kubernetes, Slurm, Volcano, Kueue, Ray, and related technologies.
  • Improve cluster utilization, reliability, elasticity, and operational efficiency.
RDMA & High-Performance Networking
  • Design and optimize RDMA, InfiniBand, RoCE, and high-speed Ethernet fabrics for distributed AI workloads.
  • Optimize GPU-to-GPU and GPU-to-NIC communication paths.
  • Improve distributed communication efficiency for large-scale training and inference.
  • Analyze and eliminate networking bottlenecks across AI clusters.
Unified KV Cache & Distributed Memory Systems
  • Design and implement unified KV Cache architecture across:
  • GPU HBM
  • CPU Memory
  • RDMA-accessible Memory
  • NVMe SSD
  • Distributed Storage
  • Develop efficient KV Cache sharing, migration, offloading, and scheduling mechanisms.
  • Optimize latency and throughput for large-scale inference systems.
CUDA & System Performance Optimisation
  • Develop and optimize CUDA kernels for training and inference workloads.
  • Profile and optimize GPU compute, memory, communication, and scheduling efficiency.
  • Contribute to low-level optimization of AI frameworks and inference engines.
  • Work on technologies such as FlashAttention, TensorRT, Triton, NCCL, CUTLASS, and custom operators.
Observability & Reliability
  • Build end-to-end observability platforms for AI infrastructure.
  • Design monitoring, logging, tracing, alerting, and troubleshooting frameworks.
  • Develop performance dashboards and SLO-driven operational systems.
  • Improve maintainability, debuggability, and operational excellence of AI platforms.
Automation & Platform Engineering
  • Build automation tools for deployment, provisioning, monitoring, and operations.
  • Develop Infrastructure-as-Code (IaC) solutions using Terraform, Ansible, and related tools.
  • Build CI/CD pipelines and engineering productivity platforms.
  • Improve platform scalability and operational efficiency.
Required Qualifications
Education
  • Bachelor's degree or above in Computer Science, Software Engineering, Electrical Engineering, or related fields.
Technical Skills
  • Strong software engineering and programming skills.
  • Excellent system design capability and strong engineering craftsmanship.
  • Strong coding standards and code quality awareness.
  • Strong sense of ownership, accountability, and execution.
System Fundamentals

Strong understanding of:

  • Operating Systems
  • Computer Networks
  • Distributed Systems
  • Data Structures and Algorithms
  • Linux Internals
Programming Languages

Proficiency in one or more of:

  • C++
  • Go
  • Python
  • Rust
AI Infrastructure Experience

Hands-on experience in one or more of:

  • Kubernetes
  • GPU Infrastructure
  • Distributed Systems
  • AI Infrastructure
  • HPC (High Performance Computing)
  • Cloud-Native Platforms
Networking Experience

Experience with:

  • RDMA
  • InfiniBand
  • RoCE/RoCEv2
  • GPUDirect
  • NCCL
  • UCX
  • High-Speed Ethernet
GPU & Performance Engineering

Experience with:

  • CUDA
  • GPU Performance Optimization
  • Multi-GPU Systems
  • Distributed Training
  • Distributed Inference
Preferred Qualifications
  • Experience building large-scale AI training or inference clusters.
  • Experience with vLLM, SGLang, TensorRT-LLM, Triton, DeepSpeed, Megatron-LM, Ray, or similar frameworks.
  • Experience with unified KV Cache systems, memory hierarchy optimisation, or distributed storage systems.
  • Experience with Kubernetes GPU Operator and NVIDIA NetworkOperator.
  • Experience with Prometheus, Grafana, Loki, OpenTelemetry, and observability platforms.
  • Experience contributing to open-source projects such as: vLLM, FlashAttention, CUTLASS, TVM, MLIR, Triton, Kubernetes, NCCL
  • Experience working across AI Infrastructure, HPC, Networking, and Silicon Systems is highly desirable.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Infrastructure Engineer
AI Infrastructure Engineer

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
AB03 - AI Infrastructure Engineer
AB03 - AI Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Engineer (ML Platform, AI Native production, Algorithm, cutting-edge technology, multinational company)
AI Infra Engineer (ML Platform, AI Native production, Algorithm, cutting-edge technology, multinational company)

DADACONSULTANTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Infrastructure Engineer | Up to $7K - 0310
AI Infrastructure Engineer | Up to $7K - 0310

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI DevOps Engineer (Cloud Infrastucture)
AI DevOps Engineer (Cloud Infrastucture)

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 57,000 - 77,000
AI Systems Infrastructure Engineer - Up to $7K - 0310
AI Systems Infrastructure Engineer - Up to $7K - 0310

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
6723 - AI Infrastructure Engineer
6723 - AI Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer - LCYL
AI Infrastructure Engineer - LCYL

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000