LLM Production & Cost-Optimized Engineer

Systems Limited

Karachi Division

On-site

PKR 3,000,000 - 5,000,000

Full time

41 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Systems Limited is seeking an experienced MLOps engineer to own production serving and scaling for LLM/agentic workloads, including inference infrastructure, load balancing, and caching. You will drive cost awareness and implement observability for failure modes.

Collaborate with GenAI Engineers on rollout strategies, model versioning, canaries, and fallbacks. You will explain token-cost dynamics to clients and help set governance for LLM workloads.

Qualifications

  • 4–6 years in platform/MLOps with hands-on LLM/GenAI production experience.
  • Strong understanding of token costs, batching, caching, and model routing.
  • Experience with LLM observability tooling (tracing, eval pipelines, version management).
  • Familiar with Azure AI Foundry, AWS Bedrock, Google Vertex AI, and open-source options like vLLM/TGI.
  • Experience creating canary/rollback strategies for probabilistic systems.
  • Able to explain token-cost changes to client stakeholders.
  • Collaborative with GenAI Engineers to balance build/run responsibilities.

Responsibilities

  • Own production serving and scaling for LLM/agentic workloads (inference infra, load balancing, caching)
  • Monitor and control inference cost — token usage, retry/loop cost, model routing decisions
  • Build observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt drift
  • Manage model/version rollout strategy (canary releases, fallback models, A/B testing)
  • Own incident response for LLM/agent production issues
  • Partner with GenAI Engineers and Agentic AI Architects on production-readiness reviews
  • Explain token-cost dynamics to client finance/business stakeholders
  • Collaborate closely with GenAI Engineers without a hard line between build and run
  • Support the practice in setting cost governance policy for LLM workloads

Skills

MLOps engineering
LLM economics
Observability tooling
Hosting platforms
Canary/rollback
Agentic workloads
Token-cost explainability
Build/run collaboration

Job description

Systems Limited is seeking an experienced MLOps engineer to own production serving and scaling for LLM/agentic workloads, including inference infrastructure, load balancing, and caching. You will drive cost awareness and implement observability for failure modes.

Collaborate with GenAI Engineers on rollout strategies, model versioning, canaries, and fallbacks. You will explain token-cost dynamics to clients and help set governance for LLM workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Production Engineer: Scaling, Cost & Observability
LLM Production Engineer: Scaling, Cost & Observability

Systems Limited • Islamabad

On-site
PKR 2,000,000 - 3,600,000
LLM Production Engineer - Cost, Rollouts & Observability
LLM Production Engineer - Cost, Rollouts & Observability

Systems Limited • Lahore

On-site
PKR 4,000,000 - 7,000,000
Forward Deployed Engineer - LLMOps
Forward Deployed Engineer - LLMOps

Systems Limited • Karachi Division

On-site
PKR 3,000,000 - 5,000,000
Forward Deployed Engineer - LLMOps
Forward Deployed Engineer - LLMOps

Systems Limited • Islamabad

On-site
PKR 2,000,000 - 3,600,000
Forward Deployed Engineer - LLMOps
Forward Deployed Engineer - LLMOps

Systems Limited • Lahore

On-site
PKR 4,000,000 - 7,000,000
Lead MLOps Engineer — Production & Observability
Lead MLOps Engineer — Production & Observability

Systems Limited • Islamabad

On-site
PKR 4,000,000 - 7,000,000
Senior AI/ML Engineer - LLM Fine-Tuning & Production
Senior AI/ML Engineer - LLM Fine-Tuning & Production

WAMO LABS • Lahore

Hybrid
PKR 300,000 - 400,000
Competitive salary
Bi-annual performance bonuses
Generous paid time off
+6
Senior MLOps Engineer: Production, FinOps & Observability
Senior MLOps Engineer: Production, FinOps & Observability

Systems Limited • Lahore

On-site
PKR 3,000,000 - 5,400,000
Senior LLM Engineer: Production AI, RAG & Agents (Remote)
Senior LLM Engineer: Production AI, RAG & Agents (Remote)

BearPlex, Inc • Lahore

On-site
PKR 1,800,000 - 4,200,000
Learning budget
Fully remote in Pakistan
AI/ML Engineer: Build Production-Ready LLM Apps
AI/ML Engineer: Build Production-Ready LLM Apps

Technology Rivers, LLC • Islamabad

On-site
PKR 3,348,000 - 4,687,000
Company-paid lunch facility
Healthcare benefits
Provident Fund (Employer Matching)
+3