Inference Systems Engineer

LambdaQ labs Pvt. Ltd.

Maharashtra

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

LambdaQ labs Pvt. Ltd. in Maharashtra is seeking a seasoned ML infra engineer to build and optimize the serving stack for fast, reliable, and affordable Vikasit Inference at India scale.

You’ll own batching, caching, routing, autoscaling, and cost management across GPU-heavy model fleets, while ensuring the OpenAI-compatible API remains rock-solid for streaming and tool calls.

Qualifications

  • Deep systems engineering experience (GPU, CUDA-adjacent, or high-performance serving).
  • Experience with an LLM serving framework in production.
  • Comfort owning latency, throughput, and cost SLOs.

Responsibilities

  • Optimize throughput and latency across model fleet (vLLM / SGLang / TensorRT-LLM).
  • Build batching, KV-cache, quantization, and routing for MoE models.
  • Own reliability, autoscaling, and cost of the inference platform.
  • Keep the OpenAI-compatible API rock-solid for streaming and tool calls.

Skills

Deep systems engineering
LLM serving framework
Latency/throughput cost SLOs

Job description

About The Role

You’ll build and optimize the serving stack that makes Vikasit Inference fast, reliable, and affordable at India scale — kernels, batching, caching, and the OpenAI-compatible API surface developers love.

What you'll do
  • Optimize throughput and latency across our model fleet (vLLM / SGLang / TensorRT-LLM)
  • Build batching, KV-cache, quantization, and routing for MoE models
  • Own reliability, autoscaling, and cost of the inference platform
  • Keep the OpenAI-compatible API rock-solid for streaming and tool calls
What we're looking for
  • Deep systems engineering (GPU, CUDA-adjacent, or high-performance serving)
  • Experience with an LLM serving framework in production
  • Comfort owning latency, throughput, and cost SLOs
Nice to have
  • CUDA / Triton kernels
  • Quantization (GGUF, AWQ, FP8)
  • Multi-region infra
Sound like you?

We hire for skill over credentials. Tell us why you're a fit — links and projects welcome.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Systems Engineer
Inference Systems Engineer

Nava • Bengaluru

On-site
INR 1,700,000 - 2,500,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Chennai District

Hybrid
INR 4,000,000 - 7,000,000
Hybrid work model
LLM Ops Engineer
LLM Ops Engineer

gnani.ai • Bengaluru

On-site
INR 2,800,000 - 4,800,000
Senior Forward Deployed Engineer I Ai Inference Digitalocean Inc Bengaluru
Senior Forward Deployed Engineer I Ai Inference Digitalocean Inc Bengaluru

Vibehackers • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Travel up to 30%
Open-source contributions
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Bengaluru

Hybrid
INR 5,500,000 - 9,000,000
Platform Engineer
Platform Engineer

Flipkart • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Senior AI Engineer
Senior AI Engineer

Arcana • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Backend / ML-Ops Engineer — Speech Model Deployment & Inference Optimization
Backend / ML-Ops Engineer — Speech Model Deployment & Inference Optimization

OutcomesAI • Bengaluru

On-site
INR 1,500,000 - 2,000,000
AI Engineer
AI Engineer

Engati Technologies Inc. • Ernakulam

On-site
INR 1,200,000 - 1,800,000
Engineering Manager - AI Engineering
Engineering Manager - AI Engineering

B Capital • Bengaluru

On-site
INR 3,800,000 - 7,000,000