Machine Learning Systems Senior Engineer

Defence Science and Technology Agency

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Defence Science and Technology Agency seeks a highly technical ML Systems Engineer to architect and build scalable AI inference capabilities across heterogeneous environments. You will work at the intersection of machine learning, systems engineering, and software engineering to create platforms and tooling that standardise and simplify AI model serving in production.

You will design abstractions for model inference, integrate across formats and environments, and optimize latency, throughput,

Qualifications

  • Degree required in Computer Science, Computer Engineering, or related field.
  • 2–3 years of relevant experience in ML systems, inference engineering, platform engineering, or performance‑critical software systems.
  • Hands‑on experience with at least one inference stack (Traditional AI, NVIDIA Triton, Generative AI, vLLM, Dynamo/LLM-D).
  • Strong ability to profile, diagnose, and optimise performance bottlenecks.
  • Proficiency in at least one programming language (e.g. Python, C++, Go, Rust).
  • Good understanding of Linux systems, distributed systems concepts, and system‑level debugging.
  • Familiarity with containers and orchestration platforms such as Docker and Kubernetes/OpenShift.

Responsibilities

  • Inference Systems Engineering: design abstractions and components to support model inference across environments.
  • Model Handling: optimize loading strategies and lifecycle management for diverse models (LLMs, CV, NLP).
  • Benchmarking & Evaluation: develop benchmarks for latency, throughput, and hardware performance.
  • Platform Integration & Developer Experience: build APIs, libraries, and services for simplified deployment and observability.

Skills

Inference stacks
Performance profiling
Programming languages
Linux & distributed
Containers & orchestration

Education

Degree in CS/CE or related

Tools

NVIDIA Triton Inference Server
vLLM
SGLang
Dynamo/LLM-D
Docker
Kubernetes/OpenShift

Job description

What the role is:

We are looking for a highly technical ML Systems Engineer to architect and build scalable AI inference capabilities across heterogeneous environments. This role focuses on solving real‑world challenges in AI model execution, runtime interoperability and performance optimization. You will operate at the intersection of machine learning, systems engineering, and software engineering, building platforms and tooling that standardise and simplify AI models serving in production environments.

What you will be working on:
  1. Inference Systems Engineering
    • Design and develop abstractions, middleware, and system components to support model inference across Traditional and Generative AI
    • Build integration layers across different model formats, execution engines, and deployment environments
    • Ensure consistency, portability, reliability, and scalability of model execution
  2. Model Handling
    • Support diverse model architectures, including: Large Language Models (LLMs), Computer vision models, NLP models, Multi‑modal models
    • Optimise models for latency, throughput and resource efficiency
    • Optimise model loading strategies
    • Implement robust mechanisms for model lifecycle management
  3. Benchmarking & Evaluation
    • Develop and execute benchmarking methodologies to evaluate: Latency vs throughput trade‑offs, Runtime and hardware performance characteristics, Use case performance characteristics
    • Support data‑driven deployment decisions through profiling and performance analysis
  4. Platform Integration & Developer Experience
    • Develop APIs, libraries, and platform services that enable: Simplified model deployment and serving, Runtime backends selection, Model Observability, Model Scaling
    • Improve developer and platform operators’ experience while preserving operational flexibility and low‑level control
What we are looking for:
Technical Experience – Must‑Have
  • Hands‑on experience with at least one inference stack: Traditional AI, NVIDIA Triton Inference Server, Generative AI, vLLM, SGLang, Dynamo/LLM-D
  • Strong ability to profile, diagnose, and optimise performance bottlenecks
  • Strong proficiency in at least one programming language (e.g. Python, C++, Go, Rust)
  • Good understanding of Linux systems, distributed systems concept, and system‑level debugging
  • Familiarity with containers and orchestration platforms such as Docker and Kubernetes/OpenShift
Preferred Experience
  • Experience working in air‑gapped or restricted environments with enterprise GPU (e.g. A100, H200, B200)
  • Experience with LLM inference, including the understanding of terms such as KV cache management, Prefill vs decode phases, continuous batching and token‑level scheduling
  • Experience with model optimisation including the understanding of terms such as quantisation (FP16, INT8, INT4), graph optimisation and compilation
Requirements
  • Degree in Computer Science, Computer Engineering, or a related discipline
  • Minimum 2–3 years of relevant experience in ML systems, inference engineering, platform engineering, or performance‑critical software systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
Senior ML Systems Engineer - Inference Platforms
Senior ML Systems Engineer - Inference Platforms

Defence Science and Technology Agency • Singapore

On-site
SGD 120,000 - 180,000
Machine Learning Engineer (Senior/Principal) - Model Training
Machine Learning Engineer (Senior/Principal) - Model Training

Michael Page • Singapore

On-site
SGD 180,000 - 260,000
Competitive salary
Bonus
GenAI exposure
+1
Machine Learning Engineer (Senior/Principal) - Model Training
Machine Learning Engineer (Senior/Principal) - Model Training

Michael Page Singapore • Singapore

On-site
SGD 120,000 - 180,000
Competitive salary
Comprehensive benefits package
Senior LLM Inference Engineer Performance & GPU Optimization
Senior LLM Inference Engineer Performance & GPU Optimization

Confidential • Singapore

On-site
SGD 90,000 - 130,000
Machine Learning Engineer (Senior/Principal) - Model Training
Machine Learning Engineer (Senior/Principal) - Model Training

michael page (personnel) pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
Competitive salary
Comprehensive benefits package
Inference Performance Engineer
Inference Performance Engineer

adaption • Singapore

Hybrid
SGD 120,000 - 160,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Ensign InfoSecurity • Singapore

On-site
SGD 120,000 - 180,000
Machine Learning Engineer
Machine Learning Engineer

Changi Airport Group • Singapore

On-site
SGD 90,000 - 150,000
Senior AI/Machine Learning Engineer
Senior AI/Machine Learning Engineer

Good co India • Singapore

On-site
SGD 120,000 - 170,000