MTS 2, AI Platform Professional

The Networker

Bengaluru

On-site

INR 3,000,000 - 5,200,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

The Networker in Bengaluru seeks an LLM Inference Engineer to remove compute-scaling bottlenecks for production LLMs, ensuring fast, reliable, observable inference from GPUs to APIs that products depend on. You’ll work at the intersection of HPC, GPU systems, and MLOps.

Join a team that optimizes runtimes (vLLM, Triton), tunes latency and throughput, builds benchmarking suites, and improves monitoring, tracing and incident response to deliver measurable production improvements.

Qualifications

  • 5+ years of development experience in AI/ML systems.
  • Experience deploying LLM inference services in production.
  • Strong Python coding skills with Go or Rust.
  • Experience with PyTorch, vLLM, TensorRT.
  • Knowledge of GPU architecture and performance tuning.
  • Exposure to quantization, batching, and caching strategies.
  • Proven ability to deliver measurable production improvements.

Responsibilities

  • Own production inference from handoff to production-grade serving.
  • Tune inference performance to reduce latency and increase throughput.
  • Optimize runtimes and servers across heterogeneous GPU fleets.
  • Benchmark and measure latency, throughput, and GPU utilization.
  • Improve monitoring, tracing, and incident response.
  • Apply optimizations like quantization and batching.
  • Collaborate with data science and product teams to meet SLOs.

Skills

Python
Go or Rust
PyTorch
vLLM
TensorRT
CUDA
Performance optimization
Root-cause analysis
Systems programming
Monitoring & observability

Tools

Triton
Benchmarking tooling

Job description

Job Summary

As an LLM Inference Engineer on our AI Platform team, you'll remove the compute-scaling bottleneck for production LLMs. Your job is to make frontier-model inference fast, efficient, reliable, and observable-the last mile from GPUs to APIs that products depend on. This role sits at the intersection of HPC, GPU systems, and MLOps, and requires strong intuition for how model architecture, runtimes, and hardware interact.


What You’ll Do

  • Own production inference: Take models from handoff to production-grade serving, including release engineering, capacity planning, cost optimization, and incident response.
  • Tune inference performance: reduce end-to-end latency and increase throughput across real production traffic patterns.
  • Optimize runtimes and servers: Scale inference across heterogeneous GPU fleets; optimize stacks such as vLLM, Triton, and related components (e.g., schedulers, KV cache, batching, memory).
  • Benchmark and measure: Build benchmarking suites, metrics, and tooling to quantify latency, throughput, GPU utilization, memory, and cost.
  • Reliability and observability: Improve monitoring, tracing, and alerting; participate in incident response and postmortems to harden systems.
  • Apply and ship new optimizations: Evaluate research and implement pragmatic inference optimizations (e.g., quantization, paging, kernel/runtimes improvements).
  • Partner with cross-functional teams: Work with data science and product teams to translate business requirements into performance and availability SLOs.

What We’re Looking For

  • 5+ years of strong development experience
  • Experience deploying and operating LLM inference services in production.
  • Strong production coding skills in Python plus Go or Rust (systems-level implementation and debugging).
  • Experience with ML frameworks and runtimes: PyTorch, vLLM, SGLang (and/or TensorRT).
  • Knowledge of GPU architecture and performance (profiling, memory bandwidth/latency tradeoffs); CUDA/kernel programming is a strong plus.
  • Solid understanding of LLM inference and optimization techniques: continuous batching, KV cache management, quantization, speculative decoding (nice-to-have), etc.
  • 3+ years hands‑on experience in performance optimization and systems programming for AI/ML workloads.
  • Demonstrated ability to deliver measurable production improvements (e.g., 2X throughput, lower p95/p99 latency, reduced GPU cost).
  • Proven skill in root‑cause analysis: finding bottlenecks across model, runtime, networking, and infrastructure.
  • Demonstrated proficiency in applying autonomous AI coding agents to speed up software delivery pipelines. This includes advanced prompting and careful human‑in‑the‑loop code review to improve development speed and code accuracy.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff
Member of Technical Staff

eBay • Bengaluru

On-site
INR 4,000,000 - 7,500,000
MTS 2, AI Platform
MTS 2, AI Platform

eBay • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Performance Engineer - Inference
Performance Engineer - Inference

Keka Inc. • Bengaluru

On-site
INR 300,000 - 540,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Bengaluru

On-site
INR 4,200,000 - 6,300,000
Inference Server Engineer
Inference Server Engineer

Evollabs • Hyderabad

On-site
INR 4,000,000 - 6,500,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Chennai District

On-site
INR 4,000,000 - 7,000,000
Hybrid work model
LLM Ops Engineer
LLM Ops Engineer

gnani.ai • Bengaluru

On-site
INR 2,800,000 - 4,800,000
Inference Systems Engineer
Inference Systems Engineer

Nava • Bengaluru

On-site
INR 1,700,000 - 2,500,000
Member of Technical Staff
Member of Technical Staff

aion • Bengaluru

On-site
INR 300,000 - 540,000
Lead MLOps Engineer
Lead MLOps Engineer

Cloudkeeper • Dadri

On-site
INR 4,500,000 - 7,000,000