Forward Deployed Engineer - LLMOps

Systems Limited

Islamabad

On-site

PKR 2,000,000 - 3,000,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Systems Limited is seeking an experienced MLOps engineer to own production serving and scaling for LLM/agentic workloads, focusing on inference infra, load balancing, and caching.

You will monitor token costs, build observability, manage canary rollouts, and collaborate with GenAI engineers on production readiness, cost governance, and incident response.

Qualifications

  • 4–6 yrs platform/MLOps engineering with hands-on LLM/GenAI production experience.
  • Deep understanding of LLM inference economics — token costs, batching, caching.
  • Experience with LLM observability tooling (tracing, eval pipelines, prompt/version management).
  • Familiarity with multiple model hosting platforms and their cost/performance tradeoffs — Azure, AWS, Vertex AI, plus self-hosted options (vLLM, TGI).
  • Experience building canary/rollback strategies for probabilistic systems.
  • Comfortable with the higher unpredictability of agentic workloads vs. classical ML serving.
  • Cost-conscious communicator — can explain token-cost dynamics to client finance stakeholders.

Responsibilities

  • Own production serving and scaling for LLM/agentic workloads (inference infra, load balancing, caching).
  • Monitor and control inference cost — token usage, retry/loop cost, model routing decisions.
  • Build observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt drift.
  • Manage model/version rollout strategy (canary releases, fallback models, A/B testing).
  • Own incident response for LLM/agent production issues.
  • Partner with GenAI Engineers and Agentic AI Architects on production-readiness reviews.
  • Explain token-cost dynamics to client finance/business stakeholders.
  • Collaborate closely with GenAI Engineers without needing a hard line between “build” and “run.”
  • Support the practice in setting cost governance policy for LLM workloads.

Skills

MLOps engineering
LLM production experience
Cost optimization
Observability tooling
Model deployment
Canary deployments
Cross-team collaboration

Tools

Azure AI Foundry
AWS Bedrock
Google Vertex AI
vLLM
TGI

Job description

ABOUT

Owns production operations for LLM and agentic workloads — serving, cost, and observability for a fundamentally less predictable class of system than classical ML.

KEY RESPONSIBILITIES
  • Own production serving and scaling for LLM/agentic workloads (inference infra, load balancing, caching)
  • Monitor and control inference cost — token usage, retry/loop cost, model routing decisions
  • Build observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt drift
  • Manage model/version rollout strategy (canary releases, fallback models, A/B testing)
  • Own incident response for LLM/agent production issues
  • Partner with GenAI Engineers and Agentic AI Architects on production-readiness reviews
  • Explain token-cost dynamics to client finance/business stakeholders
  • Collaborate closely with GenAI Engineers without needing a hard line between build and run
  • Support the practice in setting cost governance policy for LLM workloads
REQUIREMENTS & SKILLS
  • 4–6 yrs platform/MLOps engineering with hands-on LLM/GenAI production experience
  • Deep understanding of LLM inference economics — token costs, batching, caching, model routing
  • Experience with LLM observability tooling (tracing, eval pipelines, prompt/version management)
  • Familiarity with multiple model hosting platforms and their cost/performance tradeoffs — Microsoft Azure AI Foundry, AWS Bedrock, and Google Vertex AI, plus self-hosted open-source options (vLLM, TGI) as a good-to-have
  • Experience building canary/rollback strategies for probabilistic systems
  • Comfortable with the higher unpredictability of agentic workloads vs. classical ML serving
  • Cost-conscious communicator — can explain a token-cost blowup to a client's finance stakeholder
  • Collaborates closely with GenAI Engineers without needing a hard line between “build” and “run”
  • Calm under pressure during live incidents affecting client-facing systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Forward Deployed Engineer - MLOps
Forward Deployed Engineer - MLOps

Systems Limited • Islamabad

On-site
PKR 2,400,000 - 4,200,000
Senior MLOps Engineer
Senior MLOps Engineer

Systems Limited • Lahore

On-site
PKR 3,000,000 - 5,400,000
Principal Forward Deployed Engineer
Principal Forward Deployed Engineer

Systems Limited • Islamabad

On-site
PKR 3,500,000 - 5,000,000
AI Platform Engineer
AI Platform Engineer

Systems Limited • Lahore

On-site
PKR 3,000,000 - 5,400,000
AI Platform Engineer
AI Platform Engineer

Systems Limited • Islamabad

On-site
PKR 1,200,000 - 2,000,000
AI/ML Engineer
AI/ML Engineer

Glimstech • Bahawalpur Division

On-site
PKR 2,000,000 - 4,500,000
Director of Innovation Delivery
Director of Innovation Delivery

Creativechaos • Pakistan

Remote
PKR 2,000,000 - 3,000,000
Senior AI/ML Engineer
Senior AI/ML Engineer

WAMO LABS • Lahore

Hybrid
PKR 300,000 - 400,000
Competitive salary
Bi-annual performance bonuses
Generous paid time off
+6
Senior AI Engineer
Senior AI Engineer

Octdaily • Pakistan

On-site
PKR 3,600,000 - 6,000,000
AI Engineer
AI Engineer

Intelligent Inference • Islamabad

On-site
PKR 2,790,000 - 5,580,000
Equity available