Principal ML Infrastructure Engineer

PhysicsX

Singapore

On-site

SGD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity options
Private medical insurance
25 days annual leave
Pension contribution 10%
Free office lunches

Job summary

PhysicsX is seeking a Principal ML Infrastructure Engineer to design, operate, and scale the training and deployment pipelines for large physics models. You will work closely with ML researchers and engineers to ensure efficient training, robust data I/O, and reliable serving on GPU clusters.

You will champion distributed training, experiment tracking, and platform tooling, with a strong emphasis on reproducibility and performance optimization in a research-driven environment.

Qualifications

  • 5+ years of experience building and operating ML infrastructure at scale.
  • Deep expertise in distributed training and debugging NCCL hangs, with knowledge of FSDP vs. DDP vs. pipeline parallelism.
  • Strong Linux systems fundamentals, networking (NVLink/InfiniBand), storage I/O, profiling, and performance optimization.
  • Production experience with Kubernetes and SLURM for GPU cluster orchestration.
  • Proficiency in Python and ML frameworks (Python/PyTorch).
  • Experience with cloud GPU infrastructure; ideally CoreWeave or similar HPC-focused clouds.

Responsibilities

  • Design and operate distributed training infrastructure for neural operator architectures on large GPU platforms.
  • Optimize training pipelines for throughput, fault tolerance, and cost efficiency (checkpointing, gradient accumulation, multi-node sync).
  • Build and maintain experiment tracking and observability systems for training runs and model performance.

Skills

Distributed training
Linux
Networking (NVLink/InfiniBand)
Kubernetes
Python
PyTorch
SLURM
NCCL
Cloud GPU infra

Tools

NVIDIA DGX
Weights & Biases
MLflow
Prometheus
Grafana

Job description

PhysicsX is seeking a Principal ML Infrastructure Engineer to design, operate, and scale the training and deployment pipelines for large physics models. You will work closely with ML researchers and engineers to ensure efficient training, robust data I/O, and reliable serving on GPU clusters.

You will champion distributed training, experiment tracking, and platform tooling, with a strong emphasis on reproducibility and performance optimization in a research-driven environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Machine Learning Engineer Singapore
Principal Machine Learning Engineer Singapore

PhysicsX Ltd • Singapore

On-site
SGD 120,000 - 150,000
Senior Machine Learning Infrastructure Engineer, Research
Senior Machine Learning Infrastructure Engineer, Research

PhysicsX • Singapore

On-site
SGD 180,000 - 280,000
Equity options
Private medical insurance
25 days annual leave
+2
Senior AI Infra Engineer — GPU Cluster & ML Platform
Senior AI Infra Engineer — GPU Cluster & ML Platform

DADACONSULTANTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Lead ML Systems Engineer — Production-Grade AI Pipelines
Lead ML Systems Engineer — Production-Grade AI Pipelines

AI Chopping Block • Singapore

On-site
SGD 120,000 - 180,000
AI Infra Engineer (ML Platform, AI Native production, Algorithm, cutting-edge technology, multinational company)
AI Infra Engineer (ML Platform, AI Native production, Algorithm, cutting-edge technology, multinational company)

DADACONSULTANTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Software Engineer, ML Dev Enablement
Software Engineer, ML Dev Enablement

MOTIONAL SINGAPORE PTE. LIMITED • Singapore

On-site
SGD 80,000 - 130,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

K2 PARTNERING SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 180,000 - 300,000
Senior ML Infra Engineer - Multi-Cloud GPU Scaling
Senior ML Infra Engineer - Multi-Cloud GPU Scaling

Motional • Singapore

On-site
SGD 120,000 - 180,000
Senior ML Infra Engineer: Multi-Cloud & GPU Scaling
Senior ML Infra Engineer: Multi-Cloud & GPU Scaling

Motional AD Inc. • Singapore

On-site
SGD 120,000 - 180,000
Lead ML Systems Engineer: From Research to Production
Lead ML Systems Engineer: From Research to Production

Bjak • Singapore

On-site
SGD 180,000 - 260,000