Inference Engineer

Binaire Private Limited

New Delhi

On-site

INR 800,000 - 1,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI solutions company in Delhi seeks an Inference Engineer to optimize model execution and enhance inference pipelines. The ideal candidate will have strong Python skills, an understanding of machine learning inference, and experience with ML frameworks like PyTorch or TensorFlow. Responsibilities include implementing inference pipelines, optimizing performance, and collaborating with teams to enhance production features. This role offers hands-on experience with real-world AI inference at scale, along with opportunities for career growth in engineering roles.

Qualifications

  • Strong fundamentals in Python required; C++ or Rust is a plus.
  • Understanding of machine learning inference vs training.
  • Familiarity with at least one ML framework like PyTorch or TensorFlow.

Responsibilities

  • Implement and optimize model inference pipelines for LLMs and vision models.
  • Work with inference frameworks to optimize latency and throughput.
  • Assist in deploying and maintaining inference services on GPU and CPU clusters.

Skills

Python
Machine learning inference understanding
ML framework familiarity
Basic GPU/CPU architecture knowledge
Experience with Linux and Docker

Tools

TensorRT
ONNX Runtime
PyTorch
TensorFlow

Job description

We’re building a high-performance AI inference platform focused on delivering low-latency, cost-efficient model serving at scale. As an Inference Engineer, you’ll work close to the metal—optimizing model execution, improving throughput, and helping design reliable inference pipelines used in real production workloads.

This role is ideal for an engineer with strong fundamentals who wants deep exposure to model serving, hardware efficiency, and distributed systems.

What You’ll Do
  • Implement and optimize model inference pipelines for LLMs and vision models
  • Work with inference frameworks (e.g., TensorRT, ONNX Runtime, vLLM, Triton)
  • Optimize latency, throughput, memory usage, and cost per token
  • Assist in deploying and maintaining inference services on GPU and CPU clusters
  • Profile model execution (compute, memory, bandwidth) and identify bottlenecks
  • Support model quantization, batching, caching, and parallelism strategies
  • Collaborate with platform, infra, and product teams to ship production features
  • Monitor inference workloads and help improve system reliability
Required Skills
  • Strong fundamentals in Python (required); C++ or Rust is a plus
  • Understanding of machine learning inference vs training
  • Familiarity with at least one ML framework (PyTorch, TensorFlow, JAX)
  • Basic knowledge of GPU/CPU architecture, memory, and parallelism
  • Experience with Linux, containers (Docker), and basic cloud workflows
  • Ability to read research papers or performance benchmarks and apply learnings
Nice to Have
  • Exposure to LLMs (LLaMA, Mistral, Qwen, etc.) or vision models
  • Experience with quantization (INT8/FP8) or model compression
  • Familiarity with CUDA concepts, Triton kernels, or low-level optimization
  • Knowledge of distributed systems or RPC frameworks
  • Prior work on inference benchmarks or cost optimization
What You’ll Gain
  • Hands-on experience with real-world AI inference at scale
  • Deep understanding of $/token economics and system trade-offs
  • Opportunity to grow into senior inference, systems, or infra roles
  • Work on problems that directly impact product performance and margins
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Inference Engineer
Senior Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 2,600,000 - 4,800,000
Inference Systems Engineer
Inference Systems Engineer

Nava • Bengaluru

On-site
INR 1,700,000 - 2,500,000
Distributed Training & Inference Optimization Engineer
Distributed Training & Inference Optimization Engineer

Winzons • India

On-site
INR 3,000,000 - 5,000,000
AI Platform Engineer (Inference)
AI Platform Engineer (Inference)

Amazon • Bengaluru

On-site
INR 4,000,000 - 7,000,000
ML Research Engineer (Inference)
ML Research Engineer (Inference)

Cerebras Systems, Inc. • India

On-site
INR 800,000 - 1,200,000
Software Engineer II - Serverless Inference
Software Engineer II - Serverless Inference

Jobtailor • Bengaluru

On-site
INR 2,400,000 - 4,200,000
ML Research Engineer (Inference)
ML Research Engineer (Inference)

Cerebras • India

On-site
INR 5,681,000 - 9,470,000
Publish and open source cutting-edge AI research
Work on one of the fastest AI supercomputers
Job stability with startup vitality
Principal Research Engineer, Applied AI
Principal Research Engineer, Applied AI

Mulya Technologies • India

On-site
INR 1,800,000 - 3,000,000
Principal Research Engineer, Applied AI
Principal Research Engineer, Applied AI

EnCharge AI • India

On-site
INR 3,000,000 - 6,000,000
ML Research Engineer (Inference)
ML Research Engineer (Inference)

Cerebras • Bengaluru

On-site
INR 1,800,000 - 2,800,000