AI Engineer

Intelligent Inference

Islamabad

On-site

PKR 2,790,000 - 5,580,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity available

Job summary

Intelligent Inference is seeking an AI Engineer for its i2 platform to own the path from API edge to accelerator and back. You will operate the Go gateway and serving engines, deploying open-weight models on local hardware and tuning performance for enterprise deployments.

You will need strong Go production experience, hands-on LLM serving knowledge, and a solid mental model of transformer inference. Linux, containers, and Python are essential, with a focus on measurement and verifiable metrics.

Qualifications

  • Proficient Go production services engineering.
  • Hands-on experience serving LLMs in production.
  • Strong understanding of transformer inference concepts.
  • Linux, containers, and hardware topology expertise.
  • Python for model/evaluation work.
  • Ability to measure and verify performance metrics.

Responsibilities

  • Extend and operate the Go API gateway and routing logic.
  • Deploy and tune open-weight models on Ascend/NPU clusters and other stacks.
  • Own serving performance: batching, cache, parallel layouts, decode behaviour.
  • Handle quantisation work from BF16 to INT8/INT4 and assess quality impact.
  • Work within an 8-chip HCCS coherent domain per cluster and MoE routing constraints.
  • Build and operate retrieval infrastructure including vector stores and embeddings.
  • Instrument all metrics: latency, tokens/second, cost per million tokens.

Skills

Go
LLM serving
Linux
Python
Measurement/telemetry

Tools

Kubernetes
vLLM
SGLang
TensorRT-LLM
MindIE
CANN

Job description

Company Description

Intelligent Inference is building software and engineering tools that enable people to run sovereign AI inference, giving users ownership, control, and transparency over their LLM workloads. Starting with our flagship product "i2" a sovereign gateway where you can access leading open-weight AI models with agent harness and tooling, sovereignty and control.

We help developers and teams run open-source models with fast, cost-effective inference and a sovereign gateway. The company offers a full-stack, self-service AI inference platform for developers and enterprises, featuring optimized open-source model APIs with regional deployment, low latency, and reduced costs. Its platform supports dedicated inference and model deployments, fine-tuning and training, as well as bring-your-own-key and bring-your-own-model workflows. Intelligent Inference provides a comprehensive control panel with deep visibility into token usage, logs, and cost attribution, plus localized billing in home currencies without foreign exchange charges. The first sovereign regional deployment is in Pakistan, with plans to expand across the AMEA region by 2030.

AI Engineer, Inference Platform
Intelligent Inference (Private) Limited - 'i2'

i2 currently runs a commercial inference platform inside Pakistan. We self-host open-weight frontier models on Nvidia GPUs as well as Huawei Ascend NPU infrastructure, a hybrid gateway that allows you to call model APIs on both CUDA and CANN. We expose the model APIs through an OpenAI-compatible API with all native features such as streaming etc, and bill in PKR. Our customers include developers (B2C), freelancers, startups and software-houses, regulated enterprises, banks, telcos and government, who cannot legally route workloads to foreign APIs, and developers who want frontier open-weight model access in the optimal way and without USD payments. The platform is in commercial operation. Our PK model end-points are completely local, data stays resident, and the serving stack is ours end to end. That last part is what makes this an engineering job rather than a reselling job.

The role

You will own the path a request takes from the API edge to the accelerator and back. That means the Go gateway that fronts our models, and the serving engines behind it.

Concretely:
  • Extend and operate our Go API gateway: routing, authentication, rate limiting, quota enforcement, per-model metering, streaming, request shaping, failure handling.
  • Deploy and tune open-weight models on our Ascend 910B clusters using the CANN/MindIE stack, and on vLLM and SGLang where applicable.
  • Own serving performance: continuous batching, KV cache configuration, tensor and expert parallel layouts, prefill and decode behaviour, context length trade-offs.
  • Run quantisation work. Our path is BF16 to INT8 (W8A8) to INT4 (W4A16). You will run the conversions, measure the quality cost, and be honest about it.
  • Work within a hard architectural constraint: an 8-chip HCCS coherent domain per cluster. Large MoE models have to fit that shape or be sharded around it. This is the interesting part of the job.
  • Build and operate retrieval infrastructure for enterprise deployments, including vector stores, embedding endpoints and the ingestion path around them.
  • Instrument everything. Per-endpoint latency, tokens per second, time to first token, utilisation, cost per million tokens. If it is not measured we do not claim it.
  • Support dedicated enterprise deployments, including air-gapped and restricted-network environments where you cannot assume internet egress or a friendly package manager.
What we need you to already have
  • Strong Go. You have written and operated production services in it, not just read about it.
  • Real experience serving LLMs in production with vLLM, SGLang, TensorRT-LLM, TGI or equivalent. You should be able to explain why throughput fell when concurrency rose, and what you did about it.
  • A working mental model of transformer inference: attention, KV cache, batching, the prefill and decode split, why MoE routing changes the memory picture.
  • Linux, containers, and comfort at the level of drivers, kernels and hardware topology. This stack breaks below the Python line.
  • Python for model and evaluation work.
  • The habit of measuring before claiming. We mark unverified numbers as unverified, in customer documents and internally.
What will help but is not required
  • Huawei Ascend, CANN or MindIE experience. Very few people have this. If you do, say so early.
  • Quantisation experience, particularly INT8 and INT4 on non-NVIDIA silicon.
  • Kubernetes, and experience running stateful GPU or NPU workloads on it.
  • Experience with regulated deployments, audit logging, or data residency requirements.
  • Open source contributions to any inference or serving project.
What we are not asking for
  • CUDA fluency as a gate. Most of our candidates come from NVIDIA-only backgrounds and that is fine. Ascend is different silicon with a different toolchain, no FP8 or INT4 support in hardware on the 910B, and a smaller ecosystem. We will teach it. What does not transfer is patience with undocumented behaviour, so bring that.
  • Model training or research experience. This is a serving and systems role.
  • Prompt engineering or application layer work.
Why this is worth your time

There are perhaps a handful of teams in this country doing inference at the hardware level rather than wrapping someone else's API. You will work on a stack with genuine constraints, for customers with genuine compliance requirements, at a company where the serving layer is the product rather than a cost line. You will also learn an accelerator architecture that very few engineers anywhere have touched, at a point where that is becoming commercially relevant.

Compensation and terms

Depends on qualification. PKR 200,000 - 500,000 with minimum of 2 years of production level generative AI development. Equity available for the right candidate.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

10xengineers • Lahore

On-site
PKR 3,600,000 - 6,000,000
Opportunity to work on state-of-the-art AI systems
Continuous exposure to innovative AI workloads
Growth path into architectural leadership
Senior AI/ML Engineer
Senior AI/ML Engineer

WAMO LABS • Lahore

Hybrid
PKR 300,000 - 400,000
Competitive salary
Bi-annual performance bonuses
Generous paid time off
+6
AI & Backend Engineer - Production AI & LLM Systems | Pakistan
AI & Backend Engineer - Production AI & LLM Systems | Pakistan

Volga Partners • Islamabad

Hybrid
Company-sponsored medical insurance
Paid Time Off (PTO)
Holiday pay
AI Inference Platform Engineer (Go) – Open-Weight Models
AI Inference Platform Engineer (Go) – Open-Weight Models

Intelligent Inference • Islamabad

On-site
PKR 2,790,000 - 5,580,000
Equity available
AI Developer — Onsite, Sialkot Office
AI Developer — Onsite, Sialkot Office

FabTechSol • Sialkot

On-site
AI & Backend Engineer - Production AI & LLM Systems | Pakistan
AI & Backend Engineer - Production AI & LLM Systems | Pakistan

Volgapartners • Islamabad

Hybrid
Company-sponsored medical insurance
Paid Time Off (PTO)
Career growth opportunities
Senior AI Agent Engineer – Production & Orchestration
Senior AI Agent Engineer – Production & Orchestration

Smart Working • Saddar

On-site
PKR 2,500,000 - 4,500,000
Fixed Shifts: 11:30 AM - 9:00 PM PKT (
No Weekend Work
Day 1 Benefits
+2
Junior Machine Learning Engineer – AI Systems & Frameworks
Junior Machine Learning Engineer – AI Systems & Frameworks

10xengineers • Lahore

On-site
Continuous exposure to cutting-edge AI projects
Opportunity for advanced technical roles
Work with a world-class chip company
AI Agent Engineer (Remote, Full-Time) [AS311] (PK)
AI Agent Engineer (Remote, Full-Time) [AS311] (PK)

Smart Working • Saddar

On-site
PKR 2,500,000 - 4,500,000
Fixed Shifts: 11:30 AM - 9:00 PM PKT (
No Weekend Work
Day 1 Benefits
+2
AI Engineer (Generative AI)
AI Engineer (Generative AI)

Merik Solutions • Islamabad

On-site
PKR 2,000,000 - 4,000,000