Forward Deployed Engineer - LLMOps

Systems Limited

Islamabad

On-site

PKR 2,000,000 - 3,600,000

Full time

35 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Systems Limited is seeking a seasoned Platform/MLOps engineer to own the production serving and scaling of LLM/agentic workloads. You will monitor token costs, implement cost governance, and manage canary rollout strategies while enhancing observability for AI systems.

Collaborate with GenAI Engineers on production readiness, model routing, and incident response to ensure reliable AI services for clients and internal stakeholders.

Qualifications

  • 4–6 yrs platform/MLOps engineering with hands-on LLM/GenAI production experience.
  • Deep understanding of LLM inference economics—token costs, batching, caching, model routing.
  • Experience with LLM observability tooling (tracing, eval pipelines, prompt/version management).
  • Familiarity with multiple model hosting platforms and their cost/performance tradeoffs.
  • Experience building canary/rollback strategies for probabilistic systems.
  • Comfortable with the higher unpredictability of agentic workloads vs classical ML serving.
  • Cost-conscious communicator—can explain token-cost blowups to client stakeholders.

Responsibilities

  • Own production serving and scaling for LLM/agentic workloads.
  • Monitor and control inference cost—token usage, retries, model routing decisions.
  • Build observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt drift.
  • Manage model/version rollout strategy (canary releases, fallback models, A/B testing).
  • Own incident response for LLM/agent production issues.
  • Partner with GenAI Engineers and Agentic AI Architects on production-readiness reviews.
  • Explain token-cost dynamics to client finance/business stakeholders.
  • Collaborate closely with GenAI Engineers on build/run boundaries.

Skills

GenAI production
MLOps
LLM inference
Observability tooling
Model hosting platforms
Canaries/rollbacks
Cost governance
Cross-functional comms

Tools

Azure AI Foundry
AWS Bedrock
Vertex AI
vLLM
TGI

Job description

Owns production operations for LLM and agentic workloads — serving, cost, and observability for a fundamentally less predictable class of system than classical ML.

KEY RESPONSIBILITIES
  • Own production serving and scaling for LLM/agentic workloads (inference infra, load balancing, caching)
  • Monitor and control inference cost — token usage, retry/loop cost, model routing decisions
  • Build observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt drift
  • Manage model/version rollout strategy (canary releases, fallback models, A/B testing)
  • Own incident response for LLM/agent production issues
  • Partner with GenAI Engineers and Agentic AI Architects on production-readiness reviews
  • Explain token-cost dynamics to client finance/business stakeholders
  • Collaborate closely with GenAI Engineers without needing a hard line between build and run
  • Support the practice in setting cost governance policy for LLM workloads
REQUIREMENTS & SKILLS
  • 4–6 yrs platform/MLOps engineering with hands-on LLM/GenAI production experience
  • Deep understanding of LLM inference economics — token costs, batching, caching, model routing
  • Experience with LLM observability tooling (tracing, eval pipelines, prompt/version management)
  • Familiarity with multiple model hosting platforms and their cost/performance tradeoffs — Microsoft Azure AI Foundry, AWS Bedrock, and Google Vertex AI, plus self-hosted open-source options (vLLM, TGI) as a good-to-have
  • Experience building canary/rollback strategies for probabilistic systems
  • Comfortable with the higher unpredictability of agentic workloads vs. classical ML serving
  • Cost-conscious communicator — can explain a token-cost blowup to a client's finance stakeholder
  • Collaborates closely with GenAI Engineers without needing a hard line between “build” and “run”
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Forward Deployed Engineer - LLMOps
Forward Deployed Engineer - LLMOps

Systems Limited • Karachi Division

On-site
PKR 3,000,000 - 5,000,000
Forward Deployed Engineer - LLMOps
Forward Deployed Engineer - LLMOps

Systems Limited • Lahore

On-site
PKR 4,000,000 - 7,000,000
LLM Production & Cost-Optimized Engineer
LLM Production & Cost-Optimized Engineer

Systems Limited • Karachi Division

On-site
PKR 3,000,000 - 5,000,000
Forward Deployed Engineer - GenAI
Forward Deployed Engineer - GenAI

Systems Limited • Lahore

On-site
PKR 2,500,000 - 5,500,000
Forward Deployed Engineer - GenAI
Forward Deployed Engineer - GenAI

Systems Limited • Karachi Division

On-site
PKR 2,000,000 - 3,200,000
LLM Production Engineer: Scaling, Cost & Observability
LLM Production Engineer: Scaling, Cost & Observability

Systems Limited • Islamabad

On-site
PKR 2,000,000 - 3,600,000
Forward Deployed Engineer - MLOps
Forward Deployed Engineer - MLOps

Systems Limited • Lahore

On-site
PKR 3,000,000 - 5,400,000
Forward Deployed Engineer - MLOps
Forward Deployed Engineer - MLOps

Systems Limited • Islamabad

On-site
PKR 4,000,000 - 7,000,000
Forward Deployed Engineer - MLOps
Forward Deployed Engineer - MLOps

Systems Limited • Karachi Division

On-site
PKR 2,500,000 - 4,200,000
Forward Deployed Engineer - MLOps
Forward Deployed Engineer - MLOps

Systems Limited • Islamabad

On-site
PKR 2,400,000 - 4,200,000