AI Inference Platform Engineer (Go) – Open-Weight Models

Intelligent Inference

Islamabad

On-site

PKR 2,790,000 - 5,580,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity available

Job summary

Intelligent Inference is seeking an AI Engineer for its i2 platform to own the path from API edge to accelerator and back. You will operate the Go gateway and serving engines, deploying open-weight models on local hardware and tuning performance for enterprise deployments.

You will need strong Go production experience, hands-on LLM serving knowledge, and a solid mental model of transformer inference. Linux, containers, and Python are essential, with a focus on measurement and verifiable metrics.

Qualifications

  • Proficient Go production services engineering.
  • Hands-on experience serving LLMs in production.
  • Strong understanding of transformer inference concepts.
  • Linux, containers, and hardware topology expertise.
  • Python for model/evaluation work.
  • Ability to measure and verify performance metrics.

Responsibilities

  • Extend and operate the Go API gateway and routing logic.
  • Deploy and tune open-weight models on Ascend/NPU clusters and other stacks.
  • Own serving performance: batching, cache, parallel layouts, decode behaviour.
  • Handle quantisation work from BF16 to INT8/INT4 and assess quality impact.
  • Work within an 8-chip HCCS coherent domain per cluster and MoE routing constraints.
  • Build and operate retrieval infrastructure including vector stores and embeddings.
  • Instrument all metrics: latency, tokens/second, cost per million tokens.

Skills

Go
LLM serving
Linux
Python
Measurement/telemetry

Tools

Kubernetes
vLLM
SGLang
TensorRT-LLM
MindIE
CANN

Job description

Intelligent Inference is seeking an AI Engineer for its i2 platform to own the path from API edge to accelerator and back. You will operate the Go gateway and serving engines, deploying open-weight models on local hardware and tuning performance for enterprise deployments.

You will need strong Go production experience, hands-on LLM serving knowledge, and a solid mental model of transformer inference. Linux, containers, and Python are essential, with a focus on measurement and verifiable metrics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer
AI Engineer

Intelligent Inference • Islamabad

On-site
PKR 2,790,000 - 5,580,000
Equity available
Remote Go Engineer for AI Training & Evaluation
Remote Go Engineer for AI Training & Evaluation

SME Careers • Pakistan

On-site
PKR 16,703,000 - 33,408,000
Senior AI Engineer: Production ML & Model Quality
Senior AI Engineer: Production ML & Model Quality

Edge • Islamabad

On-site
PKR 1,800,000 - 3,200,000
AI/ML Engineer
AI/ML Engineer

Glimstech • Bahawalpur Division

On-site
PKR 2,000,000 - 4,500,000
AI Architect
AI Architect

Creativechaos • Pakistan

On-site
AI Platform Engineer
AI Platform Engineer

Systems Limited • Lahore

On-site
PKR 3,000,000 - 5,400,000
AI Platform Engineer
AI Platform Engineer

Systems Limited • Islamabad

On-site
PKR 1,200,000 - 2,000,000
AI Engineer
AI Engineer

Synares • Lahore

Remote
PKR 1,800,000 - 3,200,000
Competitive salary with equity options
Comprehensive health, dental, and vision insurance
100% remote work environment
+5
Senior AI/ML Engineer
Senior AI/ML Engineer

Archisurance • Lahore

On-site
PKR 16,759,000 - 22,347,000
AI Platform Engineer — SaaS, LLMs & MLOps
AI Platform Engineer — SaaS, LLMs & MLOps

Scalemill • Karachi Division

On-site
PKR 1,800,000 - 3,200,000