Performance Engineer - Inference

Keka Inc.

Bengaluru

On-site

INR 300,000 - 540,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Keka Inc. in Bengaluru is seeking a Performance Engineer focused on inference to understand and improve systems that serve our foundation models. You will measure end-to-end performance across model execution, serving runtimes, and distributed components.

Collaborate with model, platform, and infrastructure teams to identify bottlenecks, build profiling and observability tools, and land high-impact optimizations while preserving numerical correctness.

Qualifications

  • Hands-on profiling of ML systems or other performance-critical production systems.
  • Production experience with at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, or an equivalent serving runtime.
  • Experience serving or operating large models across accelerator-backed infrastructure, including multi-GPU or multi-accelerator systems.
  • Strong Python skills and the ability to read, instrument, and modify large production codebases.

Responsibilities

  • Run cross-layer performance investigations across throughput, latency, memory efficiency, reliability, and cost.
  • Build profiling, benchmarking, and observability tools that make inference performance measurable and explainable.
  • Identify bottlenecks across model servers, batching and scheduling, distributed execution, memory systems, and accelerators.
  • Partner with model, platform, and infrastructure teams to prioritize and land high-impact optimizations.
  • Validate that performance improvements preserve model quality and numerical correctness.

Skills

Performance profiling
Python
Distributed systems
Transformer inference

Tools

vLLM
TensorRT-LLM
NVIDIA Dynamo
Profiling tools

Job description

We are looking for a Performance Engineer, Inference to understand and improve the systems that

serve our foundation models. Inference is a tightly coupled system spanning model execution, serving

runtimes, distributed systems, accelerators, scheduling, memory, and reliability. You will measure the

system end to end, identify the highest-leverage performance gaps, and work across teams to close

them while preserving correctness.

Responsibilities
  • Run cross-layer performance investigations across throughput, latency, memory efficiency, reliability, and cost.
  • Build profiling, benchmarking, and observability tools that make inference performance measurable and explainable.
  • Identify bottlenecks across model servers, batching and scheduling, distributed execution, memory systems, and accelerators.
  • Partner with model, platform, and infrastructure teams to prioritize and land high-impact optimizations.
  • Validate that performance improvements preserve model quality and numerical correctness.
Minimum Qualifications
  • Hands-on experience profiling and optimizing ML systems or other performance-critical production systems.
  • Production experience with at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, or an equivalent serving runtime.
  • Experience serving or operating large models across accelerator-backed infrastructure, including multi-GPU or multi-accelerator systems.
  • Strong Python skills and the ability to read, instrument, and modify large production codebases.
  • Solid understanding of transformer inference, distributed systems, latency/throughput trade-offs, and accelerator performance fundamentals.
Preferred Qualifications
  • Experience with large-scale or multi-node inference, including tensor, pipeline, data, or expert parallelism.
  • Experience with GPUs, TPUs, NPUs, or other ML accelerators and associated profiling tools.
  • Experience with quantization, low-precision inference, KV-cache optimization, speculative decoding, or long-context serving.
  • Experience contributing to or modifying inference runtimes, kernels, compilers, or distributed serving components.
  • Experience optimizing inference for constrained or on-device environments.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Chennai District

On-site
INR 4,000,000 - 7,000,000
Hybrid work model
MTS 2, AI Platform Professional
MTS 2, AI Platform Professional

The Networker • Bengaluru

On-site
INR 3,000,000 - 5,200,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Bengaluru

On-site
INR 4,200,000 - 6,300,000
Member of Technical Staff
Member of Technical Staff

eBay • Bengaluru

On-site
INR 4,000,000 - 7,500,000
Inference Systems Engineer
Inference Systems Engineer

Nava • Bengaluru

On-site
INR 1,700,000 - 2,500,000
Inference Server Engineer
Inference Server Engineer

Evollabs • Hyderabad

On-site
INR 4,000,000 - 6,500,000
Principal Research Engineer, Applied AI
Principal Research Engineer, Applied AI

Mulya Technologies • India

On-site
INR 4,500,000 - 9,000,000
Principal Research Engineer, Applied AI
Principal Research Engineer, Applied AI

EnCharge AI • India

On-site
INR 400,000 - 700,000
ML Systems Performance Engineer
ML Systems Performance Engineer

Cerebras Systems, Inc. • India

On-site
INR 1,200,000 - 1,800,000
LLM Ops Engineer
LLM Ops Engineer

gnani.ai • Bengaluru

On-site
INR 2,800,000 - 4,800,000