Senior Inference Engineer

Binaire Private Limited

New Delhi

On-site

INR 2,600,000 - 4,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology company in New Delhi is seeking a Senior Inference Engineer to drive architectural decisions in high-performance AI inference systems. This role entails designing end-to-end architectures for LLM and multimodal models, leading optimization efforts, and ensuring production-grade reliability across diverse workloads. The ideal candidate will have strong experience in Python and C++, deep knowledge of ML inference internals, and hands-on expertise with GPU architecture. This is a unique opportunity to work on critical AI systems and influence technical direction.

Qualifications

  • Strong experience in Python and C++ (or Rust) for performance-critical systems.
  • Deep understanding of ML inference internals (transformers, attention, KV cache).
  • Experience running production inference at scale (multi-node, multi-GPU systems).
  • Strong knowledge of GPU architecture and memory hierarchies.
  • Experience with multi-node, multi-GPU deployments and tooling.

Responsibilities

  • Design and own end-to-end inference architecture for LLM and multimodal models.
  • Lead optimization of latency, throughput, and cost per token.
  • Profile and debug performance across compute, memory, and networking.
  • Build and scale inference services across heterogeneous hardware (GPU/CPU/accelerators).
  • Evaluate and integrate inference engines (TensorRT-LLM, vLLM, Triton, custom runtimes).
  • Profile and debug performance across compute, memory, interconnect, and networking.
  • Establish production best practices: autoscaling, rollout strategies, monitoring, SLOs.
  • Mentor engineers and review performance-critical code.
  • Partner with product and business teams to translate requirements into system design.

Skills

Python
C++
ML inference internals
GPU architecture
CUDA
GPU architecture
CUDA
Triton
Linux containers
Production scale

Tools

Triton
CUDA

Job description

We’re building a high-performance AI inference platform designed for massive scale, low latency, and industry-leading cost efficiency. As a Senior Inference Engineer, you will own critical parts of the inference stack—driving architectural decisions, pushing hardware efficiency limits, and ensuring production-grade reliability across diverse workloads.

This role is for engineers who thrive at the intersection of systems engineering, ML performance, and infrastructure economics.

What You’ll Do
  • Design and own end-to-end inference architecture for LLM and multimodal models
  • Lead optimization of latency, throughput, tail latency (p95/p99), and cost per token
  • Architect batching, KV-cache management, speculative decoding, and parallelism strategies
  • Build and scale inference services across heterogeneous hardware (GPU, CPU, accelerators)
  • Evaluate and integrate inference engines (TensorRT-LLM, vLLM, Triton, custom runtimes)
  • Profile and debug performance across compute, memory, interconnect, and networking
  • Establish production best practices: autoscaling, rollout strategies, monitoring, SLOs
  • Mentor engineers and review performance-critical code
  • Partner with product and business teams to translate requirements into system design
Required Skills
  • Strong experience in Python and C++ (or Rust) for performance-critical systems
  • Deep understanding of ML inference internals (transformers, attention, KV cache)
  • Proven experience optimizing inference for LLMs or large vision models
  • Strong knowledge of GPU architecture, memory hierarchies, and parallel programming
  • Hands-on experience with CUDA, Triton, or similar kernel-level optimization tools
  • Experience running production inference at scale (multi-node, multi-GPU systems)
  • Solid background in Linux, containers, and cloud / bare-metal infrastructure
Nice to Have
  • Experience with custom CUDA kernels or compiler toolchains
  • Familiarity with distributed inference (tensor/pipeline parallelism, NCCL)
  • Knowledge of inference on alternative hardware (Trainium, Inferentia, TPUs, ASICs)
  • Experience optimizing for power efficiency and $/token economics
  • Contributions to open-source inference frameworks or performance tooling
What You’ll Gain
  • Ownership of core inference systems used in production at scale
  • Direct impact on cost structure, margins, and customer experience
  • Opportunity to shape technical direction and platform architecture
  • Work on some of the most performance-critical systems in applied AI
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Engineer
Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 800,000 - 1,200,000
Inference Performance Engineer
Inference Performance Engineer

adaption • India

On-site
INR 12,440,000 - 18,182,000
Flexible work: In-person collaboration
Adaption Passport: Annual travel
Lunch Stipend
+1
AI Inference Engineer – LLM
AI Inference Engineer – LLM

GyanSys Inc. • Bengaluru

On-site
INR 1,000,000 - 1,600,000
Distributed Training & Inference Optimization Engineer
Distributed Training & Inference Optimization Engineer

Winzons • India

On-site
INR 3,000,000 - 5,000,000
Inference Systems Engineer
Inference Systems Engineer

Nava • Bengaluru

On-site
INR 1,700,000 - 2,500,000
Principal Research Engineer, Applied AI
Principal Research Engineer, Applied AI

EnCharge AI • India

On-site
INR 3,000,000 - 6,000,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Bengaluru

Hybrid
INR 5,500,000 - 9,000,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Chennai District

Hybrid
INR 4,000,000 - 7,000,000
Hybrid work model
Senior System Software Engineer - LocalAI
Senior System Software Engineer - LocalAI

NVIDIA Gruppe • Pune District

On-site
INR 300,000 - 550,000
Senior System Software Engineer - Local AI
Senior System Software Engineer - Local AI

NVIDIA Corporation • Pune District

On-site
INR 2,500,000 - 5,000,000