Lead MLOps & SRE Engineer — Developer Experience

HTX (Home Team Science & Technology Agency)

Singapore

On-site

SGD 120,000 - 180,000

Full time

11 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

HTX (Home Team Science & Technology Agency) is seeking a seasoned MLOps/SRE/DevOps specialist to deploy and operate production-grade LLMs on GPU infrastructure. You will manage vector databases, orchestration services, and API gateways, while tuning performance and ensuring reliable, observable systems.

Applicants should bring 4+ years in ML Ops and strong Kubernetes expertise, plus IaC experience and familiarity with vector DBs, Ray/Kubeflow/MLflow. Singapore-based role with a two-year contract.

Qualifications

  • 4+ years of experience in MLOps, SRE, or DevOps roles, with at least 1 year working with ML/AI systems.
  • Hands-on experience deploying and operating LLMs in production (vLLM, TGI, TensorRT-LLM, or similar).
  • Strong Kubernetes expertise including operators, StatefulSets, and GPU scheduling.
  • Deep understanding of GPU architecture, and inference optimization techniques.
  • Experience with observability tools (Prometheus, Grafana, ELK/Elastic Stack).
  • Solid Python and Bash scripting skills for automation.
  • Knowledge of vector databases (Milvus, Weaviate, Qdrant, or Pinecone).
  • Experience with infrastructure-as-code (Terraform, Helm, Kustomize).
  • Experience with NVIDIA GPUs (A100/H100/B200) and DCGM monitoring.
  • Understanding of LLM inference concepts: KV cache, continuous batching, PagedAttention.
  • Familiarity with Ray clusters, Kubeflow, or MLflow.
  • Background in SRE practices: SLO/SLI definition, error budgets, incident management.
  • Experience with secure or regulated environments.
  • Knowledge of LiteLLM, Kong Gateway, or API management platforms.
  • Systems thinking with ability to diagnose complex issues across the ML stack.
  • Data-driven decision making using metrics and telemetry.
  • Proactive mindset focused on reliability, automation, and preventive measures.
  • Strong debugging skills for GPU, networking, and distributed systems issues.
  • Clear incident communication and documentation.
  • Collaborative approach working with data scientists, ML engineers, and platform teams.

Responsibilities

  • LLM Deployment: Deploy and manage LLM models using vLLM/TensorRT-LLM on our GPU infrastructure, optimizing for throughput, latency, and GPU utilization.
  • Infrastructure Management: Provision and maintain supporting infrastructure including vector databases, orchestration services, Redis/queue systems, and API gateways.
  • Performance Optimization: Profile and tune LLM inference performance, experiment with batching, context caching, and quantization techniques.
  • Observability: Implement monitoring using Prometheus, Grafana, DCGM exporters, and Elastic Stack to track latency, throughput, cache hits, and health.
  • Reliability Engineering: Establish SLOs/SLIs, implement auto-scaling, design failure recovery, and conduct chaos engineering for uptime.
  • Cost Optimization: Monitor GPU utilization and inference costs, identify opportunities, reduce token usage and compute spend.
  • Security & Compliance: Ensure components operate within secure boundaries, manage secrets, and maintain audit logs.
  • Incident Response: Participate in on-call rotation, troubleshoot incidents, perform root cause analysis, implement preventive measures.
  • Capacity Planning: Model future load, forecast GPU requirements, and coordinate with infra teams to scale.

Skills

Kubernetes
Python
Bash
GPU-architecture
Observability
SRE-practices
Terraform
Helm
Kustomize
DCGM
Vector-databases
Ray Kubeflow MLflow

Tools

Milvus
Weaviate
Qdrant
Pinecone
vLLM
TensorRT-LLM
TGI
Kubeflow
MLflow
NVIDIA GPUs
DCGM monitoring

Job description

HTX (Home Team Science & Technology Agency) is seeking a seasoned MLOps/SRE/DevOps specialist to deploy and operate production-grade LLMs on GPU infrastructure. You will manage vector databases, orchestration services, and API gateways, while tuning performance and ensuring reliable, observable systems.

Applicants should bring 4+ years in ML Ops and strong Kubernetes expertise, plus IaC experience and familiarity with vector DBs, Ray/Kubeflow/MLflow. Singapore-based role with a two-year contract.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior MLOps Engineer: LLM Production & SRE
Senior MLOps Engineer: LLM Production & SRE

Home Team Science and Technology Agency (HTX) • Singapore

On-site
SGD 120,000 - 170,000
Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud
Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud

HTX (Home Team Science & Technology Agency) • Singapore

On-site
SGD 120,000 - 180,000
Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud
Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud

Home Team Science and Technology Agency (HTX) • Singapore

On-site
SGD 120,000 - 170,000
Geospatial AI Engineer — ML Systems & MLOps Lead
Geospatial AI Engineer — ML Systems & MLOps Lead

TALENTSIS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior LLMOps Engineer: Production AI Reliability Lead
Senior LLMOps Engineer: Production AI Reliability Lead

NCS Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
Health insurance
Learning & development
Lead ML & AI Systems Engineer
Lead ML & AI Systems Engineer

GMP RECRUITMENT SERVICES (S) PTE LTD • Singapore

On-site
SGD 120,000 - 180,000
MLOps Engineer
MLOps Engineer

Epergne Solutions • Singapore

On-site
SGD 120,000 - 180,000
Senior MLOps Engineer: Build Production-Scale ML Pipelines
Senior MLOps Engineer: Build Production-Scale ML Pipelines

EPAM Systems • Singapore

On-site
SGD 120,000 - 180,000
MLOps Platform Engineer: AI Pipelines & LLMs
MLOps Platform Engineer: AI Pipelines & LLMs

STRT.ASIA PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Senior AI Engineer: ML Systems & MLOps Leader
Senior AI Engineer: ML Systems & MLOps Leader

TALENTSIS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000