Get more replies from employers
Send a job-specific resume in minutes.
Intelligent Inference is seeking an AI Engineer for its i2 platform to own the path from API edge to accelerator and back. You will operate the Go gateway and serving engines, deploying open-weight models on local hardware and tuning performance for enterprise deployments.
You will need strong Go production experience, hands-on LLM serving knowledge, and a solid mental model of transformer inference. Linux, containers, and Python are essential, with a focus on measurement and verifiable metrics.
Intelligent Inference is building software and engineering tools that enable people to run sovereign AI inference, giving users ownership, control, and transparency over their LLM workloads. Starting with our flagship product "i2" a sovereign gateway where you can access leading open-weight AI models with agent harness and tooling, sovereignty and control.
We help developers and teams run open-source models with fast, cost-effective inference and a sovereign gateway. The company offers a full-stack, self-service AI inference platform for developers and enterprises, featuring optimized open-source model APIs with regional deployment, low latency, and reduced costs. Its platform supports dedicated inference and model deployments, fine-tuning and training, as well as bring-your-own-key and bring-your-own-model workflows. Intelligent Inference provides a comprehensive control panel with deep visibility into token usage, logs, and cost attribution, plus localized billing in home currencies without foreign exchange charges. The first sovereign regional deployment is in Pakistan, with plans to expand across the AMEA region by 2030.
i2 currently runs a commercial inference platform inside Pakistan. We self-host open-weight frontier models on Nvidia GPUs as well as Huawei Ascend NPU infrastructure, a hybrid gateway that allows you to call model APIs on both CUDA and CANN. We expose the model APIs through an OpenAI-compatible API with all native features such as streaming etc, and bill in PKR. Our customers include developers (B2C), freelancers, startups and software-houses, regulated enterprises, banks, telcos and government, who cannot legally route workloads to foreign APIs, and developers who want frontier open-weight model access in the optimal way and without USD payments. The platform is in commercial operation. Our PK model end-points are completely local, data stays resident, and the serving stack is ours end to end. That last part is what makes this an engineering job rather than a reselling job.
You will own the path a request takes from the API edge to the accelerator and back. That means the Go gateway that fronts our models, and the serving engines behind it.
There are perhaps a handful of teams in this country doing inference at the hardware level rather than wrapping someone else's API. You will work on a stack with genuine constraints, for customers with genuine compliance requirements, at a company where the serving layer is the product rather than a cost line. You will also learn an accelerator architecture that very few engineers anywhere have touched, at a point where that is becoming commercially relevant.
Depends on qualification. PKR 200,000 - 500,000 with minimum of 2 years of production level generative AI development. Equity available for the right candidate.