Platform Engineer

Flipkart

Bengaluru

On-site

INR 4,000,000 - 6,000,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Flipkart is hiring a Senior Systems / Platform Engineer (SDE 3) in Bengaluru to build and optimize high-throughput AI model inference infrastructure. You will focus on core model execution, serving platforms, and hardware acceleration to deliver sub-second responses across search surfaces.

Join a team that emphasizes low-latency systems, memory optimization, and GPU/NPU acceleration, working with vLLM, Triton, TensorRT, and related tooling in a scale-driven environment.

Qualifications

  • Must have hands-on experience in C++ or high-performance Java/Rust with strong system design knowledge.
  • Must have direct experience with inference servers and frameworks such as vLLM, Triton Inference Server, TensorRT, ONNX Runtime, or CUDA.
  • Nice-to-have background in embedding models into vector search engines or kernel optimization for NPUs/GPUs.

Responsibilities

  • Architect and optimize high-throughput, low-latency AI model inference engines for sub-second latency at scale.
  • Tune memory layouts, KV-cache execution, batching, and model quantization (FP8, AWQ).
  • Integrate AI model inference into vector search and high-scale platform surfaces.

Skills

C++
Java
Rust

Tools

vLLM
Triton Inference Server
TensorRT
ONNX Runtime
CUDA

Job description

Flipkart is committed to the cause of transforming commerce in India through our investments in

made-in-India technology innovations, customer-centric features and constructs, a diverse category landscape and a world-class supply chain. With a customer base of over 350 million, product coverage of over 150 million across 80+ categories, focus on generating direct and indirect employment and a commitment to empowering generations of entrepreneurs and MSMEs and a sustainable growth strategy – Flipkart is maximizing for our customers, stakeholders, and the planet at large! Flipkart is a part of the Walmart-owned Flipkart Group, which also includes group companies Flipkart Health+, Myntra, and Cleartrip. The Group is also a majority shareholder in PhonePe, one of the leading Payments Apps in India.

About the Team:

You will join the Core Search & AI Platform team responsible for driving intelligent discovery across high-scale surfaces like Search, Minutes, Shopsy, and Cleartrip. We build low-latency, high-throughput model serving infrastructure and deep learning pipelines that execute sub-second inference at massive scale. Our focus is on systems engineering, memory optimization, and hardware acceleration (GPU/NPU)—pushing the boundaries of C++, CUDA, and specialized inference engines to power real-time AI across the entire ecosystem.

About the role:

We are seeking a Senior Systems / Platform Engineer (SDE 3) to build and optimize our high-throughput, low-latency AI Model Inference Infrastructure. In this role, you will focus on core model execution, serving platforms, and hardware acceleration to deliver sub-second response times across high-scale search and platform surfaces.

What You’ll Work On:
  • Architecting and tuning high-performance LLM model serving engines for sub-second latency under high QPS.
  • Optimizing memory layout, KV-cache execution, continuous batching, and model quantization (FP8, AWQ).
  • Embedding AI model inference directly into low-latency vector and hybrid search platforms.
What We Are Looking For:
  • Must-Have: Strong hands-on C++ or high-performance Java/Rust engineering experience with deep system design fundamentals.
  • Must-Have: Direct experience with inference servers and frameworks such as vLLM, Triton Inference Server, TensorRT, ONNX Runtime, or CUDA.
  • Nice-to-Have: Background in embedding models into vector search engines (FAISS, Milvus, ScaNN) or custom NPU/GPU kernel optimization.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Systems Engineer
Inference Systems Engineer

LambdaQ labs Pvt. Ltd. • Maharashtra

On-site
INR 1,200,000 - 1,800,000
Engineering Manager - AI Engineering
Engineering Manager - AI Engineering

B Capital • Bengaluru

On-site
INR 3,800,000 - 7,000,000
AI-ML Engineer
AI-ML Engineer

KanthamAi • Mumbai

On-site
INR 2,000,000 - 3,000,000
Principal Machine Learning Engineer
Principal Machine Learning Engineer

SourcingXPress • Hyderabad

On-site
INR 1,906,000 - 2,860,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Chennai District

Hybrid
INR 4,000,000 - 7,000,000
Hybrid work model
Senior/Principal Local Llm & Generative Ai Platform Engineer
Senior/Principal Local Llm & Generative Ai Platform Engineer

Parallelwireless • Maharashtra

On-site
INR 3,000,000 - 5,500,000
Senior Forward Deployed Engineer I Ai Inference Digitalocean Inc Bengaluru
Senior Forward Deployed Engineer I Ai Inference Digitalocean Inc Bengaluru

Vibehackers • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Travel up to 30%
Open-source contributions
Senior Al Engineer | Java, LLM & Gen AI
Senior Al Engineer | Java, LLM & Gen AI

Cornerstone • Hyderabad

Hybrid
INR 2,500,000 - 5,000,000
Senior AI Engineer
Senior AI Engineer

Flipkart • Bengaluru Urban

On-site
INR 1,500,000 - 2,500,000
Senior Software Engineer(AI/ML Platform)
Senior Software Engineer(AI/ML Platform)

Autodesk • Pune District

On-site
INR 1,500,000 - 2,000,000