Inference Systems Engineer

Nava

Bengaluru

On-site

INR 1,700,000 - 2,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nava is seeking an AI/ML deployment engineer to design and deploy low-latency inference pipelines for LLMs, diffusion models, and vision transformers. You will optimize serving stacks across CPU, GPU, and NPU with Triton Inference Server and ONNX Runtime, and you will containerize services using Docker and Kubernetes to ensure high availability and auto-scaling.

You will implement monitoring, health checks, and A/B testing to validate performance and drift in production, and collaborate with ML

Qualifications

  • Proficient in Python and ML model deployment.
  • Experience with PyTorch, ONNX Runtime and TensorRT.
  • Hands-on with containerization and CI/CD pipelines.

Responsibilities

  • Design and deploy low-latency inference pipelines for multiple models across CPU, GPU, and NPU.
  • Optimize serving stacks with Triton Inference Server, vLLM, or TGI for peak performance.
  • Containerize services with Docker and orchestrate with Kubernetes for HA and scalability.
  • Implement monitoring, health checks, and A/B testing for production models.
  • Collaborate with ML engineers to productionize models with minimal accuracy loss.
  • Build observability dashboards (Prometheus/Grafana) and alerting for bottlenecks.

Skills

Python
PyTorch
ONNX Runtime
TensorRT
CI/CD (GitHub Actions, GitLab CI)

Tools

Docker
Kubernetes
Triton Inference Server
Prometheus
Grafana

Job description

Role & Responsibilities
  • Design and deploy low-latency inference pipelines for LLMs, diffusion models, and vision transformers across CPU, GPU, and NPU architectures.
  • Optimize model serving stacks using Triton Inference Server, vLLM, or TGI—tuning batch sizes, quantization, and memory layout for peak performance.
  • Containerize and orchestrate inference services via Docker and Kubernetes, ensuring high availability and auto-scaling under fluctuating workloads.
  • Implement model monitoring, health checks, and A/B testing frameworks to validate performance and drift in production.
  • Collaborate with ML Engineers to convert trained models into production-ready formats (ONNX, TensorRT, GGUF) with minimal accuracy loss.
  • Build observability dashboards (Prometheus/Grafana) and alerting logic to detect and mitigate inference bottlenecks in real time.
Skills & Qualifications
Must-Have
  • Python
  • Docker
  • Kubernetes
  • Triton Inference Server
  • ONNX Runtime
  • PyTorch
  • TensorRT
  • Prometheus
  • Grafana
  • CI/CD (GitHub Actions, GitLab CI)
Preferred
  • Experience with vLLM or TGI
  • Knowledge of Model Quantization (AWQ, GPTQ, GGUF)
  • Familiarity with NVIDIA Triton model ensemble pipelines
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Engineer
Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 800,000 - 1,200,000
Senior Inference Engineer
Senior Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 2,600,000 - 4,800,000
AI Platform Engineer (Inference)
AI Platform Engineer (Inference)

Amazon • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior ML/AI Engineer
Senior ML/AI Engineer

Syren Cloud Inc. • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Distributed Training & Inference Optimization Engineer
Distributed Training & Inference Optimization Engineer

Winzons • India

On-site
INR 3,000,000 - 5,000,000
SDE II - ML/AI Engineer
SDE II - ML/AI Engineer

Aurigait • Jaipur

On-site
INR 1,200,000 - 1,800,000
Backend / ML-Ops Engineer — Speech Model Deployment & Inference Optimization
Backend / ML-Ops Engineer — Speech Model Deployment & Inference Optimization

OutcomesAI • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Bengaluru

On-site
INR 3,500,000 - 7,500,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Bengaluru

Hybrid
INR 5,500,000 - 9,000,000
ML Research Engineer (Inference)
ML Research Engineer (Inference)

Cerebras Systems, Inc. • India

On-site
INR 800,000 - 1,200,000