An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Flipkart is hiring a Senior Systems / Platform Engineer (SDE 3) in Bengaluru to build and optimize high-throughput AI model inference infrastructure. You will focus on core model execution, serving platforms, and hardware acceleration to deliver sub-second responses across search surfaces.
Join a team that emphasizes low-latency systems, memory optimization, and GPU/NPU acceleration, working with vLLM, Triton, TensorRT, and related tooling in a scale-driven environment.
Flipkart is committed to the cause of transforming commerce in India through our investments in
made-in-India technology innovations, customer-centric features and constructs, a diverse category landscape and a world-class supply chain. With a customer base of over 350 million, product coverage of over 150 million across 80+ categories, focus on generating direct and indirect employment and a commitment to empowering generations of entrepreneurs and MSMEs and a sustainable growth strategy – Flipkart is maximizing for our customers, stakeholders, and the planet at large! Flipkart is a part of the Walmart-owned Flipkart Group, which also includes group companies Flipkart Health+, Myntra, and Cleartrip. The Group is also a majority shareholder in PhonePe, one of the leading Payments Apps in India.
You will join the Core Search & AI Platform team responsible for driving intelligent discovery across high-scale surfaces like Search, Minutes, Shopsy, and Cleartrip. We build low-latency, high-throughput model serving infrastructure and deep learning pipelines that execute sub-second inference at massive scale. Our focus is on systems engineering, memory optimization, and hardware acceleration (GPU/NPU)—pushing the boundaries of C++, CUDA, and specialized inference engines to power real-time AI across the entire ecosystem.
We are seeking a Senior Systems / Platform Engineer (SDE 3) to build and optimize our high-throughput, low-latency AI Model Inference Infrastructure. In this role, you will focus on core model execution, serving platforms, and hardware acceleration to deliver sub-second response times across high-scale search and platform surfaces.