Principal Machine Learning Engineer

Visa Hunt

Singapore

On-site

SGD 180,000 - 260,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Our client, a stealth AI startup, is seeking a Principal Machine Learning Engineer to build and deploy production-grade ML systems powering its AI platform. You will translate research into scalable solutions by developing robust training pipelines, inference systems, evaluation frameworks, and deployment infrastructure.

Working closely with research and application engineering teams, you will deliver reliable, high-performance ML systems under production constraints such as latency, cost,

Qualifications

  • Strong background in deep learning and transformer-based architectures.
  • Hands-on experience training, fine-tuning, or deploying large-scale ML models in production.
  • Proficiency with PyTorch or JAX.
  • Experience with distributed training and inference frameworks (DeepSpeed, FSDP, Megatron, ZeRO, Ray).
  • Strong software engineering skills and production-grade system design.
  • Experience optimizing GPU workloads (memory, quantization, mixed precision).
  • Ability to own end-to-end ML systems in fast-moving environments.

Responsibilities

  • Build and own end-to-end ML pipelines (data processing, training, evaluation, inference, deployment).
  • Fine-tune and adapt models using LoRA, QLoRA, SFT, DPO, and distillation.
  • Design scalable inference systems balancing latency, cost, and reliability.
  • Develop data pipelines for synthetic and real-world datasets.
  • Build evaluation frameworks for performance, robustness, safety, and bias.
  • Optimize deployments with GPU optimization and scaling strategies.
  • Collaborate with backend, mobile, and desktop teams to integrate ML systems.
  • Iterate rapidly and monitor real-world performance under production constraints.

Skills

Deep learning
Transformer architectures
Production ML
Software engineering
GPU optimization
Problem solving

Tools

PyTorch
JAX
DeepSpeed
FSDP
Megatron
Ray

Job description

About the Company

Our client is a stealth AI startup backed by one of Southeast Asia's leading technology companies and is currently building its global founding team.

The company is developing an AI-native communication platform designed to simplify everyday tasks by integrating AI directly into conversations. Instead of switching between multiple applications, users can plan, organize, compare, research, and complete tasks within a single intelligent assistant.

Serving a market of billions of users still relying on traditional productivity tools, the platform focuses on delivering reliable AI workflows, persistent context, multi-step reasoning, and seamless task execution. The mission is to create an AI assistant that significantly improves productivity while making everyday work simpler and more intuitive.

About the Role

Our client is seeking a Principal Machine Learning Engineer to build and deploy production-grade machine learning systems that power its AI platform. This role focuses on translating research into scalable solutions by developing robust training pipelines, inference systems, evaluation frameworks, and deployment infrastructure.

Working closely with research and application engineering teams, this position will play a key role in delivering reliable, high-performance ML systems that operate effectively under real-world production constraints.

Key Responsibilities
  • Build and own end-to-end machine learning pipelines covering data processing, model training, evaluation, inference, and deployment.
  • Fine-tune and adapt models using modern techniques such as LoRA, QLoRA, Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and model distillation.
  • Design and operate scalable inference systems while balancing latency, cost, and reliability.
  • Develop and maintain data pipelines for both synthetic and real-world training datasets.
  • Build evaluation frameworks to assess model performance, robustness, safety, and bias in collaboration with research teams.
  • Optimize production deployments through GPU optimization, memory efficiency, latency reduction, and scaling strategies.
  • Collaborate with application engineering teams to integrate machine learning systems into backend, mobile, and desktop applications.
  • Continuously improve ML systems through rapid iteration and real-world performance monitoring while balancing production constraints such as latency, cost, reliability, and safety.
Requirements
  • Strong background in deep learning and transformer-based architectures.
  • Hands‑on experience training, fine‑tuning, or deploying large‑scale machine learning models in production.
  • Proficiency with modern machine learning frameworks such as PyTorch or JAX.
  • Experience with distributed training and inference frameworks, including technologies such as DeepSpeed, FSDP, Megatron, ZeRO, or Ray.
  • Strong software engineering skills with experience building robust, maintainable, production‑grade systems.
  • Experience optimizing GPU workloads, including memory efficiency, quantization, and mixed precision.
  • Ability to independently own end-to-end machine learning systems in fast-moving environments.
  • Strong problem‑solving skills with a focus on rapid iteration and continuous improvement.
Preferred Qualifications

Experience with one or more of the following is preferred:

  • LLM inference frameworks such as vLLM, TensorRT-LLM, or FasterTransformer
  • Open‑source contributions to machine learning or systems libraries
  • Scientific computing, compiler technologies, or GPU kernel development
  • Reinforcement Learning from Human Feedback (RLHF) pipelines, including PPO, DPO, or ORPO
  • Training or deploying multimodal or diffusion models
  • Large-scale data processing frameworks such as Apache Arrow, Spark, or Ray

Originally posted on Himalayas

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. AI/ML Engineer
Sr. AI/ML Engineer

GECO Asia Pte Ltd • Singapore

Hybrid
SGD 150,000 - 210,000
Machine Learning Engineer
Machine Learning Engineer

Changi Airport Group • Singapore

On-site
SGD 90,000 - 150,000
Machine Learning Engineer
Machine Learning Engineer

HCLTech • Singapore

On-site
SGD 90,000 - 130,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

K2 PARTNERING SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 180,000 - 300,000
Machine Learning / AI Engineer
Machine Learning / AI Engineer

GMP RECRUITMENT SERVICES (S) PTE LTD • Singapore

On-site
SGD 120,000 - 180,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Ensign InfoSecurity • Singapore

On-site
SGD 120,000 - 180,000
Lead Machine Engineer
Lead Machine Engineer

GRABTAXI HOLDINGS PTE. LTD. • Singapore

On-site
SGD 180,000 - 260,000
Sr. AI/ML Engineer
Sr. AI/ML Engineer

Tap Growth ai • Singapore

On-site
SGD 140,000 - 200,000
AI Engineer (Permanent)
AI Engineer (Permanent)

TALENTSIS PTE. LTD. • Singapore

On-site
SGD 120,000 - 160,000
Machine Learning Engineer (AI Infrastructure)
Machine Learning Engineer (AI Infrastructure)

JAC RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000