LLM Production Engineer - Cost, Scale & Observability

Systems Limited

Malaysia

On-site

MYR 120,000 - 240,000

Full time

24 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Systems Limited in Malaysia is seeking an experienced MLOps Engineer to own production serving and scaling for LLM/agentic workloads, focusing on inference infrastructure, load balancing, and caching.

You will monitor costs, build observability for failure modes, manage model rollout and incident response, and work with GenAI Engineers on production readiness and cost governance. This role requires strong collaboration with cross-functional teams and a keen eye for efficiency.

Qualifications

  • 4–6 yrs platform/MLOps engineering with hands-on LLM/GenAI production experience.
  • Deep understanding of LLM inference economics — token costs, batching, caching, model routing.
  • Experience with LLM observability tooling (tracing, eval pipelines, prompt/version management).
  • Familiarity with multiple model hosting platforms and cost/performance tradeoffs — Azure AI Foundry, AWS Bedrock, Vertex AI, and self-hosted options (vLLM, TGI).
  • Experience building canary/rollback strategies for probabilistic systems.
  • Comfortable with higher unpredictability of agentic workloads vs classical ML serving.
  • Cost-conscious communicator — explain token-cost dynamics to client finance stakeholders.
  • Collaborates closely with GenAI Engineers without a hard line between build and run.

Responsibilities

  • Own production serving and scaling for LLM/agentic workloads (inference infra, load balancing, caching).
  • Monitor and control inference cost — token usage, retry/loop cost, model routing decisions.
  • Build observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt drift.
  • Manage model/version rollout strategy (canary releases, fallback models, A/B testing).
  • Own incident response for LLM/agent production issues.
  • Partner with GenAI Engineers and Agentic AI Architects on production-readiness reviews.
  • Explain token-cost dynamics to client finance/business stakeholders.
  • Collaborate closely with GenAI Engineers without needing a hard line between build and run.
  • Support the practice in setting cost governance policy for LLM workloads.

Skills

MLOps engineering
LLM inference economics
Observability tooling
Model hosting platforms
Canary/rollback strategies
Cost-conscious communication
GenAI collaboration
Open-source options (vLLM, TGI)

Tools

Azure AI Foundry
AWS Bedrock
Google Vertex AI
vLLM
TGI

Job description

Systems Limited in Malaysia is seeking an experienced MLOps Engineer to own production serving and scaling for LLM/agentic workloads, focusing on inference infrastructure, load balancing, and caching.

You will monitor costs, build observability for failure modes, manage model rollout and incident response, and work with GenAI Engineers on production readiness and cost governance. This role requires strong collaboration with cross-functional teams and a keen eye for efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Forward Deployed Engineer - LLMOps
Forward Deployed Engineer - LLMOps

Systems Limited • Malaysia

On-site
MYR 120,000 - 240,000
Production-Ready AI Solutions Engineer (LLMs & Automation)
Production-Ready AI Solutions Engineer (LLMs & Automation)

techstreet • Petaling Jaya

On-site
MYR 60,000 - 110,000
GenAI Engineer: LLMs, RAG & MLOps in Production
GenAI Engineer: LLMs, RAG & MLOps in Production

Amast • Batu 6 ½ Jalan Puchong

On-site
MYR 80,000 - 120,000
ML & AI Engineer: Production Systems & MLOps
ML & AI Engineer: Production Systems & MLOps

Maxis • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Senior MLOps & Production Reliability Engineer
Senior MLOps & Production Reliability Engineer

Systems Limited • Kuala Lumpur

On-site
MYR 180,000 - 240,000
AI Solutions Engineer: Production-Ready LLMs & Automation
AI Solutions Engineer: Production-Ready LLMs & Automation

techstreet • Selangor

On-site
MYR 90,000 - 130,000
Senior MLOps Engineer – AI Production & Cloud
Senior MLOps Engineer – AI Production & Cloud

DKSH Malaysia Sdn Bhd • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Senior MLOps Analyst — Production ML Engineer
Senior MLOps Analyst — Production ML Engineer

Nabla • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Hybrid work
Flexible working environment
Volunteer time off
+2
GenAI Engineer: LLMs, MLOps & AI Innovation Asia
GenAI Engineer: LLMs, MLOps & AI Innovation Asia

QBE Insurance • Petaling Jaya

On-site
MYR 120,000 - 180,000
GenAI Engineer (LLMs, MLOps) - 12-Month Contract
GenAI Engineer (LLMs, MLOps) - 12-Month Contract

QBE Insurance Group • Petaling Jaya

On-site
MYR 100,000 - 180,000