Inference Systems Engineer

Nava

Bengaluru

On-site

INR 1,700,000 - 2,500,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Nava is seeking an AI/ML deployment engineer to design and deploy low-latency inference pipelines for LLMs, diffusion models, and vision transformers. You will optimize serving stacks across CPU, GPU, and NPU with Triton Inference Server and ONNX Runtime, and you will containerize services using Docker and Kubernetes to ensure high availability and auto-scaling.

You will implement monitoring, health checks, and A/B testing to validate performance and drift in production, and collaborate with ML

Qualifications

  • Proficient in Python and ML model deployment.
  • Experience with PyTorch, ONNX Runtime and TensorRT.
  • Hands-on with containerization and CI/CD pipelines.

Responsibilities

  • Design and deploy low-latency inference pipelines for multiple models across CPU, GPU, and NPU.
  • Optimize serving stacks with Triton Inference Server, vLLM, or TGI for peak performance.
  • Containerize services with Docker and orchestrate with Kubernetes for HA and scalability.
  • Implement monitoring, health checks, and A/B testing for production models.
  • Collaborate with ML engineers to productionize models with minimal accuracy loss.
  • Build observability dashboards (Prometheus/Grafana) and alerting for bottlenecks.

Skills

Python
PyTorch
ONNX Runtime
TensorRT
CI/CD (GitHub Actions, GitLab CI)

Tools

Docker
Kubernetes
Triton Inference Server
Prometheus
Grafana

Job description

Role & Responsibilities
  • Design and deploy low-latency inference pipelines for LLMs, diffusion models, and vision transformers across CPU, GPU, and NPU architectures.
  • Optimize model serving stacks using Triton Inference Server, vLLM, or TGI—tuning batch sizes, quantization, and memory layout for peak performance.
  • Containerize and orchestrate inference services via Docker and Kubernetes, ensuring high availability and auto-scaling under fluctuating workloads.
  • Implement model monitoring, health checks, and A/B testing frameworks to validate performance and drift in production.
  • Collaborate with ML Engineers to convert trained models into production-ready formats (ONNX, TensorRT, GGUF) with minimal accuracy loss.
  • Build observability dashboards (Prometheus/Grafana) and alerting logic to detect and mitigate inference bottlenecks in real time.
Skills & Qualifications
Must-Have
  • Python
  • Docker
  • Kubernetes
  • Triton Inference Server
  • ONNX Runtime
  • PyTorch
  • TensorRT
  • Prometheus
  • Grafana
  • CI/CD (GitHub Actions, GitLab CI)
Preferred
  • Experience with vLLM or TGI
  • Knowledge of Model Quantization (AWQ, GPTQ, GGUF)
  • Familiarity with NVIDIA Triton model ensemble pipelines
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff
Member of Technical Staff

eBay • Bengaluru

On-site
INR 4,000,000 - 7,500,000
Performance Engineer - Inference
Performance Engineer - Inference

Keka Inc. • Bengaluru

On-site
INR 300,000 - 540,000
MTS 2, AI Platform Professional
MTS 2, AI Platform Professional

The Networker • Bengaluru

On-site
INR 3,000,000 - 5,200,000
Inference Server Engineer
Inference Server Engineer

Evollabs • Hyderabad

On-site
INR 4,000,000 - 6,500,000
LLM Ops Engineer
LLM Ops Engineer

gnani.ai • Bengaluru

On-site
INR 2,800,000 - 4,800,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Chennai District

On-site
INR 4,000,000 - 7,000,000
Hybrid work model
SDE II - ML/AI Engineer
SDE II - ML/AI Engineer

Aurigait • Jaipur

On-site
INR 1,200,000 - 1,800,000
AI Engineer Model Optimization & Acceleration
AI Engineer Model Optimization & Acceleration

Sunrise Biztech Systems • Bangalore Rural

On-site
INR 1,200,000 - 2,400,000
AI ML Ops Engineer
AI ML Ops Engineer

Keka Inc. • Indore District

On-site
INR 1,800,000 - 3,000,000
Senior Forward Deployed Engineer I Ai Inference Digitalocean Inc Bengaluru
Senior Forward Deployed Engineer I Ai Inference Digitalocean Inc Bengaluru

Vibehackers • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Travel up to 30%
Open-source contributions