Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud

HTX (Home Team Science & Technology Agency)

Singapore

On-site

SGD 120,000 - 180,000

Full time

8 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

HTX (Home Team Science & Technology Agency) is seeking a seasoned MLOps/SRE/DevOps specialist to deploy and operate production-grade LLMs on GPU infrastructure. You will manage vector databases, orchestration services, and API gateways, while tuning performance and ensuring reliable, observable systems.

Applicants should bring 4+ years in ML Ops and strong Kubernetes expertise, plus IaC experience and familiarity with vector DBs, Ray/Kubeflow/MLflow. Singapore-based role with a two-year contract.

Qualifications

  • 4+ years of experience in MLOps, SRE, or DevOps roles, with at least 1 year working with ML/AI systems.
  • Hands-on experience deploying and operating LLMs in production (vLLM, TGI, TensorRT-LLM, or similar).
  • Strong Kubernetes expertise including operators, StatefulSets, and GPU scheduling.
  • Deep understanding of GPU architecture, and inference optimization techniques.
  • Experience with observability tools (Prometheus, Grafana, ELK/Elastic Stack).
  • Solid Python and Bash scripting skills for automation.
  • Knowledge of vector databases (Milvus, Weaviate, Qdrant, or Pinecone).
  • Experience with infrastructure-as-code (Terraform, Helm, Kustomize).
  • Experience with NVIDIA GPUs (A100/H100/B200) and DCGM monitoring.
  • Understanding of LLM inference concepts: KV cache, continuous batching, PagedAttention.
  • Familiarity with Ray clusters, Kubeflow, or MLflow.
  • Background in SRE practices: SLO/SLI definition, error budgets, incident management.
  • Experience with secure or regulated environments.
  • Knowledge of LiteLLM, Kong Gateway, or API management platforms.
  • Systems thinking with ability to diagnose complex issues across the ML stack.
  • Data-driven decision making using metrics and telemetry.
  • Proactive mindset focused on reliability, automation, and preventive measures.
  • Strong debugging skills for GPU, networking, and distributed systems issues.
  • Clear incident communication and documentation.
  • Collaborative approach working with data scientists, ML engineers, and platform teams.

Responsibilities

  • LLM Deployment: Deploy and manage LLM models using vLLM/TensorRT-LLM on our GPU infrastructure, optimizing for throughput, latency, and GPU utilization.
  • Infrastructure Management: Provision and maintain supporting infrastructure including vector databases, orchestration services, Redis/queue systems, and API gateways.
  • Performance Optimization: Profile and tune LLM inference performance, experiment with batching, context caching, and quantization techniques.
  • Observability: Implement monitoring using Prometheus, Grafana, DCGM exporters, and Elastic Stack to track latency, throughput, cache hits, and health.
  • Reliability Engineering: Establish SLOs/SLIs, implement auto-scaling, design failure recovery, and conduct chaos engineering for uptime.
  • Cost Optimization: Monitor GPU utilization and inference costs, identify opportunities, reduce token usage and compute spend.
  • Security & Compliance: Ensure components operate within secure boundaries, manage secrets, and maintain audit logs.
  • Incident Response: Participate in on-call rotation, troubleshoot incidents, perform root cause analysis, implement preventive measures.
  • Capacity Planning: Model future load, forecast GPU requirements, and coordinate with infra teams to scale.

Skills

Kubernetes
Python
Bash
GPU-architecture
Observability
SRE-practices
Terraform
Helm
Kustomize
DCGM
Vector-databases
Ray Kubeflow MLflow

Tools

Milvus
Weaviate
Qdrant
Pinecone
vLLM
TensorRT-LLM
TGI
Kubeflow
MLflow
NVIDIA GPUs
DCGM monitoring

Job description

What The Role Is

HTX is Singapore's Science and Technology agency that brings together diverse scientific and engineering capabilities to develop transformative, operationally ready solutions for public safety. As a statutory board under the Ministry of Home Affairs, HTX works at the forefront of science and technology to empower the Home Team with cutting-edge capabilities. Guided by our mission to amplify, augment and accelerate the Home Team's advantage, we are committed to keeping Singapore the safest place on planet Earth.

xCloud is dedicated to elevating the enterprise experience through cutting-edge cloud capabilities, including:
  • Sustainable enterprise data centres
  • Hybrid cloud platforms
  • Advanced cloud security measures
  • Integrated Development, Security, and Operations (DevSecOps)
  • Innovative enterprise Software-as-a-Service (SaaS) solutions
  • AI-powered enterprise applications
What You Will Be Working On
  • LLM Deployment: Deploy and manage LLM models using vLLM/TensorRT-LLM on our GPU infrastructure, optimizing for throughput, latency, and GPU utilization
  • Infrastructure Management: Provision and maintain the supporting infrastructure including vector databases (for RAG), orchestration services, Redis/queue systems, and API gateways
  • Performance Optimization: Profile and tune LLM inference performance, experiment with batching strategies, context caching, and quantization techniques to maximize throughput within GPU constraints
  • Observability: Implement comprehensive monitoring using Prometheus, Grafana, DCGM exporters, and Elastic Stack to track inference latency, token throughput, cache hit rates, and system health
  • Reliability Engineering: Establish SLOs/SLIs, implement auto-scaling policies, design failure recovery mechanisms, and conduct chaos engineering to ensure high uptime
  • Cost Optimization: Monitor GPU utilization and inference costs, identify optimization opportunities, and implement strategies to reduce token usage and compute spend
  • Security & Compliance: Ensure all components operate within secure network boundaries, manage secrets and credentials securely, and maintain audit logs for compliance
  • Incident Response: Participate in on-call rotation, troubleshoot production incidents, conduct root cause analysis, and implement preventive measures
  • Capacity Planning: Model future load, forecast GPU requirements, and work with infrastructure teams to scale the platform as adoption grows
What We Are Looking For
  • 4+ years of experience in MLOps, SRE, or DevOps roles, with at least 1 year working with ML/AI systems
  • Hands‑on experience deploying and operating LLMs in production (vLLM, TGI, TensorRT‑LLM, or similar)
  • Strong Kubernetes expertise including operators, StatefulSets, and GPU scheduling
  • Deep understanding of GPU architecture, and inference optimization techniques
  • Experience with observability tools (Prometheus, Grafana, ELK/Elastic Stack)
  • Solid Python and Bash scripting skills for automation
  • Knowledge of vector databases (Milvus, Weaviate, Qdrant, or Pinecone)
  • Experience with infrastructure‑as‑code (Terraform, Helm, Kustomize)
  • Experience with NVIDIA GPUs (A100/H100/B200) and DCGM monitoring
  • Understanding of LLM inference concepts: KV cache, continuous batching, PagedAttention
  • Familiarity with Ray clusters, Kubeflow, or MLflow
  • Background in SRE practices: SLO/SLI definition, error budgets, incident management
  • Experience with secure or regulated environments
  • Knowledge of LiteLLM, Kong Gateway, or API management platforms
  • Systems thinking with ability to diagnose complex issues across the ML stack
  • Data‑driven decision making using metrics and telemetry
  • Proactive mindset focused on reliability, automation, and preventive measures
  • Strong debugging skills for GPU, networking, and distributed systems issues
  • Clear incident communication and documentation
  • Collaborative approach working with data scientists, ML engineers, and platform teams

All new hires are appointed on a two-year contract in the first instance and will be assessed and considered for permanent tenure over time, based on performance.

As part of the shortlisting process for this role, you may be required to complete a medical declaration and/or undergo further assessment.

All applicants will be updated on the status of their applications within 4 weeks upon closing of the advertisement.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud
Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud

Home Team Science and Technology Agency (HTX) • Singapore

On-site
SGD 120,000 - 170,000
Lead MLOps & SRE Engineer — Developer Experience
Lead MLOps & SRE Engineer — Developer Experience

HTX (Home Team Science & Technology Agency) • Singapore

On-site
SGD 120,000 - 180,000
Deputy Director, AI R&D (LLM & Multimodal AI), Q Team CoE
Deputy Director, AI R&D (LLM & Multimodal AI), Q Team CoE

HTX (Home Team Science & Technology Agency) • Singapore

On-site
SGD 180,000 - 240,000
Deputy Director, AI R&D (LLM & Multimodal AI), Q Team CoE
Deputy Director, AI R&D (LLM & Multimodal AI), Q Team CoE

Home Team Science and Technology Agency (HTX) • Singapore

On-site
SGD 120,000 - 160,000
AI Engineer, Platforms
AI Engineer, Platforms

Nanyang Technological University Singapore • Singapore

On-site
SGD 90,000 - 140,000
Senior AI Engineer (Xora Portfolio Company)
Senior AI Engineer (Xora Portfolio Company)

Xora Innovation • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Full-Stack Engineer — LLMs & Cloud
Senior AI Full-Stack Engineer — LLMs & Cloud

KEYSIGHT TECHNOLOGIES SINGAPORE (SALES) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Member of Technical Staff, Machine Learning
Member of Technical Staff, Machine Learning

Alphasearch • Singapore

On-site
SGD 120,000 - 180,000
Head, MLOPs, AI Platform, xCloud
Head, MLOPs, AI Platform, xCloud

Home Team Science and Technology Agency (HTX) • Singapore

On-site
SGD 200,000 - 260,000
Sr. AI/ML Engineer
Sr. AI/ML Engineer

GECO Asia Pte Ltd • Singapore

Hybrid
SGD 150,000 - 210,000