Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud

Home Team Science and Technology Agency (HTX)

Singapore

On-site

SGD 75,000 - 100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Home Team Science and Technology Agency (HTX) in Singapore is seeking an MLOps/SRE Engineer to optimize production LLM systems. You will manage the full stack from LLM inference to observability, ensuring high reliability and performance.

The successful candidate will have strong experience in MLOps, Kubernetes management, and GPU architecture, contributing to advanced AI infrastructure aimed at enhancing Singapore’s homeland security.

Qualifications

  • 4+ years in MLOps, SRE, or DevOps roles, with 1+ year with ML/AI systems.
  • Hands-on experience deploying LLMs in production.
  • Strong Kubernetes expertise including operators and GPU scheduling.
  • Deep understanding of GPU architecture and inference optimization.

Responsibilities

  • Deploy and manage LLM models on GPU infrastructure.
  • Provision and maintain supporting infrastructure including vector databases.
  • Profile and tune LLM inference performance for maximum throughput.
  • Implement comprehensive monitoring using observability tools.

Skills

MLOps
Site Reliability Engineering (SRE)
DevOps
Kubernetes
Python scripting
Bash scripting
GPU architecture
Observability tools

Tools

Prometheus
Grafana
Elastic Stack
Terraform
Helm
Kustomize

Job description

What the role is

HTX is the first Science and Technology Agency of its kind in the world, bringing together science and engineering capabilities across the Home Team Departments to transform Singapore’s homeland security landscape. We are a statutory board under the Ministry of Home Affairs, dedicated to developing cutting‑edge technologies that empower our Home Team to solve crimes, save lives, secure borders, and safeguard public spaces. As the MLOps/SRE Engineer for HTX's developer experience squad, you will be responsible for deploying, operating, and optimizing a production LLM system in our secure infrastructure. You will ensure the agentic code assistant is reliable, performant, and cost‑effective, managing the full stack from LLM inference to vector databases, orchestration services, and observability. This role combines deep MLOps expertise with SRE discipline to support Home Team’s critical AI infrastructure.

What you will be working on
  • LLM Deployment: Deploy and manage LLM models using vLLM/TensorRT‑LLM on our GPU infrastructure, optimizing for throughput, latency, and GPU utilization
  • Infrastructure Management: Provision and maintain the supporting infrastructure including vector databases (for RAG), orchestration services, Redis/queue systems, and API gateways
  • Performance Optimization: Profile and tune LLM inference performance, experiment with batching strategies, context caching, and quantization techniques to maximize throughput within GPU constraints
  • Observability: Implement comprehensive monitoring using Prometheus, Grafana, DCGM exporters, and Elastic Stack to track inference latency, token throughput, cache hit rates, and system health
  • Reliability Engineering: Establish SLOs/SLIs, implement auto‑scaling policies, design failure recovery mechanisms, and conduct chaos engineering to ensure high uptime
  • Cost Optimization: Monitor GPU utilization and inference costs, identify optimization opportunities, and implement strategies to reduce token usage and compute spend
  • Security & Compliance: Ensure all components operate within secure network boundaries, manage secrets and credentials securely, and maintain audit logs for compliance
  • Incident Response: Participate in on‑call rotation, troubleshoot production incidents, conduct root cause analysis, and implement preventive measures
  • Capacity Planning: Model future load, forecast GPU requirements, and work with infrastructure teams to scale the platform as adoption grows
What we are looking for
  • 4+ years of experience in MLOps, SRE, or DevOps roles, with at least 1 year working with ML/AI systems
  • Hands‑on experience deploying and operating LLMs in production (vLLM, TGI, TensorRT‑LLM, or similar)
  • Strong Kubernetes expertise including operators, StatefulSets, and GPU scheduling
  • Deep understanding of GPU architecture, and inference optimization techniques
  • Experience with observability tools (Prometheus, Grafana, ELK/Elastic Stack)
  • Solid Python and Bash scripting skills for automation
  • Knowledge of vector databases (Milvus, Weaviate, Qdrant, or Pinecone)
  • Experience with infrastructure‑as‑code (Terraform, Helm, Kustomize)
  • Experience with NVIDIA GPUs (A100/H100/B200) and DCGM monitoring
  • Understanding of LLM inference concepts: KV cache, continuous batching, PagedAttention
  • Familiarity with Ray clusters, Kubeflow, or MLflow
  • Background in SRE practices: SLO/SLI definition, error budgets, incident management
  • Experience with secure or regulated environments
  • Knowledge of LiteLLM, Kong Gateway, or API management platforms

Competencies:

  • Systems thinking with ability to diagnose complex issues across the ML stack
  • Data‑driven decision making using metrics and telemetry
  • Proactive mindset focused on reliability, automation, and preventive measures
  • Strong debugging skills for GPU, networking, and distributed systems issues
  • Clear incident communication and documentation
  • Collaborative approach working with data scientists, ML engineers, and platform teams

All new hires are appointed on a two‑year contract in the first instance and will be assessed and considered for permanent tenure over time, based on performance.

About Home Team Science and Technology Agency (HTX)

HTX is the world’s first Science and Technology agency that integrates a diverse range of scientific and engineering capabilities to innovate and deliver transformative and operationally‑ready solutions for homeland security. As a statutory board of the Ministry of Home Affairs and integral to the Home Team, HTX works at the forefront of science and technology to empower Singapore’s frontline of security. Our shared mission is to amplify, augment and accelerate the Home Team’s advantage and secure Singapore as the safest place on planet earth.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Engineer / Engineer, AI R&D (LLM), Q Team CoE
Lead Engineer / Engineer, AI R&D (LLM), Q Team CoE

Home Team Science and Technology Agency (HTX) • Singapore

On-site
SGD 70,000 - 100,000
Conference attendance
Continuous learning opportunities
Lead Engineer/ Engineer, DevOps & MLOps (AI Platform), xCloud
Lead Engineer/ Engineer, DevOps & MLOps (AI Platform), xCloud

Home Team Science and Technology Agency (HTX) • Singapore

On-site
SGD 90,000 - 130,000
Head, MLOPs, AI Platform, xCloud
Head, MLOPs, AI Platform, xCloud

Home Team Science and Technology Agency (HTX) • Singapore

On-site
SGD 120,000 - 160,000
Head, MLOPs, AI Platform, xCloud
Head, MLOPs, AI Platform, xCloud

HTX (Home Team Science & Technology Agency) • Singapore

On-site
SGD 180,000 - 260,000
Two-year contract
Career progression
Professional development
Lead Engineer/ Engineer, DevOps & MLOps (AI Platform), xCloud
Lead Engineer/ Engineer, DevOps & MLOps (AI Platform), xCloud

HTX (Home Team Science & Technology Agency) • Singapore

On-site
SGD 90,000 - 150,000
Head, Cloud Platform Engineering, Cloud Engineering, xCloud
Head, Cloud Platform Engineering, Cloud Engineering, xCloud

Home Team Science and Technology Agency (HTX) • Singapore

On-site
SGD 120,000 - 160,000
Deputy Director, AI R&D (LLM & Multimodal AI), Q Team CoE
Deputy Director, AI R&D (LLM & Multimodal AI), Q Team CoE

HTX (Home Team Science & Technology Agency) • Singapore

On-site
SGD 300,000 - 420,000
Lead Engineer/Engineer, AI Safety and Security, AI R&D, xCyber
Lead Engineer/Engineer, AI Safety and Security, AI R&D, xCyber

HTX (Home Team Science & Technology Agency) • Singapore

On-site
SGD 70,000 - 90,000
Lead Engineer/Engineer, AI Safety and Security, AI R&D, xCyber
Lead Engineer/Engineer, AI Safety and Security, AI R&D, xCyber

Home Team Science and Technology Agency (HTX) • Singapore

On-site
SGD 70,000 - 90,000
Lead Engineer/Engineer, AI & Data Engineering, xData
Lead Engineer/Engineer, AI & Data Engineering, xData

Home Team Science and Technology Agency (HTX) • Singapore

On-site
SGD 70,000 - 100,000