AI Inference Engineer – LLM

GyanSys Inc.

Bengaluru

On-site

INR 1,000,000 - 1,600,000

Full time

3 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GyanSys Inc. is seeking an AI Inference Engineer in Bengaluru to own end-to-end inference pipelines, performance tuning, and scalable production serving for GenAI workloads.

You will optimize models, runtimes, and deployments using vLLM/TensorRT-LLM, CUDA, and Kubernetes, while collaborating with ML researchers on training and fine-tuning to improve inference efficiency.

Qualifications

  • 5-7 years of experience in AI/ML engineering, inference systems, GPU computing or distributed systems.
  • Strong programming experience in Python; C++/CUDA experience is a strong advantage.
  • Strong understanding of Transformer-based models and LLM architectures.
  • Hands-on experience with PyTorch and Hugging Face Transformers.
  • Practical experience with one or more inference frameworks such as vLLM, SGLang or TensorRT-LLM.

Responsibilities

  • Design, develop and optimize LLM/SLM and Generative AI inference pipelines for production workloads.
  • Deploy and scale models using inference frameworks such as vLLM, SGLang, TensorRT-LLM or equivalent.
  • Optimize latency, throughput, GPU utilization, memory footprint and inference cost.
  • Work on KV-cache optimization, continuous batching, speculative decoding, quantization and model parallelism.
  • Optimize inference across multi-GPU and multi-node environments.
  • Analyze GPU compute and memory bottlenecks using profiling and benchmarking tools.
  • Work with CUDA, NCCL, GPU memory management and NVIDIA GPU architectures.
  • Evaluate and benchmark models across different GPU configurations and inference runtimes.
  • Develop model serving solutions using Kubernetes, Docker and GPU orchestration platforms.
  • Collaborate with ML engineers and researchers on model training, fine‑tuning and inference optimization.
  • Support SFT, LoRA/QLoRA and PEFT workflows and understand their impact on inference performance.
  • Build automated performance benchmarking and evaluation frameworks for AI models.
  • Implement monitoring and observability for production inference workloads.
  • Work with MLOps/LLMOps teams on model versioning, deployment, CI/CD and production lifecycle management.
  • Troubleshoot production issues involving GPU utilization, memory fragmentation, latency, throughput and scaling.

Skills

Python
C++/CUDA
Transformer models
PyTorch
Hugging Face
vLLM / TensorRT-LLM
Kubernetes
Docker
MLOps
GPGPU/Performance

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
NCCL

Job description

We are looking for an AI Inference Engineer with strong hands‑on experience in LLM/GenAI model inference, optimization, deployment and GPU computing. The engineer will work across the AI inference stack, from model optimization and runtime development to scalable production serving.

The ideal candidate should have a strong understanding of Transformer architectures, GPU systems, inference runtimes and distributed computing, with experience optimizing AI workloads for performance, latency, throughput and cost.

Key Responsibilities
  • Design, develop and optimize LLM/SLM and Generative AI inference pipelines for production workloads.
  • Deploy and scale models using inference frameworks such as vLLM, SGLang, TensorRT-LLM or equivalent.
  • Optimize latency, throughput, GPU utilization, memory footprint and inference cost.
  • Work on KV-cache optimization, continuous batching, speculative decoding, quantization and model parallelism.
  • Optimize inference across multi-GPU and multi-node environments.
  • Analyze GPU compute and memory bottlenecks using profiling and benchmarking tools.
  • Work with CUDA, NCCL, GPU memory management and NVIDIA GPU architectures.
  • Evaluate and benchmark models across different GPU configurations and inference runtimes.
  • Develop model serving solutions using Kubernetes, Docker and GPU orchestration platforms.
  • Collaborate with ML engineers and researchers on model training, fine‑tuning and inference optimization.
  • Support SFT, LoRA/QLoRA and PEFT workflows and understand their impact on inference performance.
  • Build automated performance benchmarking and evaluation frameworks for AI models.
  • Implement monitoring and observability for production inference workloads.
  • Work with MLOps/LLMOps teams on model versioning, deployment, CI/CD and production lifecycle management.
  • Troubleshoot production issues involving GPU utilization, memory fragmentation, latency, throughput and scaling.
Required Skills
  • 5-7 years of experience in AI/ML engineering, inference systems, GPU computing or distributed systems.
  • Strong programming experience in Python; C++/CUDA experience is a strong advantage.
  • Strong understanding of Transformer‑based models and LLM architectures.
  • Hands‑on experience with PyTorch and Hugging Face Transformers.
  • Practical experience with one or more inference frameworks such as vLLM, SGLang or TensorRT-LLM.
  • Strong understanding of GPU architecture, CUDA, GPU memory and NCCL.
  • Experience with multi‑GPU and distributed AI workloads.
  • Understanding of KV cache, batching, quantization, attention optimization and parallelism.
  • Experience with Docker and Kubernetes for deploying AI workloads.
  • Strong experience in performance benchmarking and optimization.
Good to Have
  • Experience with NVIDIA H100/H200/A100/B200 or newer GPU architectures.
  • Experience with TensorRT, CUDA kernels or CUDA Graphs.
  • Experience with DeepSpeed, Megatron‑LM, FSDP or Ray.
  • Experience with LLM fine‑tuning, SFT, LoRA/QLoRA or DPO.
  • Experience with MLflow, LLMOps or MLOps platforms.
  • Knowledge of GPU networking, NVLink, InfiniBand and high‑performance computing.
  • Experience with inference observability, profiling and cost optimization.
What You'll Work On

The role spans the complete AI model‑to‑production inference stack:

The successful candidate should be able to understand both the AI model and the underlying compute infrastructure, and make informed trade‑offs between latency, throughput, memory, scalability and cost.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Engineer
Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 800,000 - 1,200,000
Senior ML/AI Engineer
Senior ML/AI Engineer

Syren Cloud Inc. • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Senior Inference Engineer
Senior Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 2,600,000 - 4,800,000
AI Architect
AI Architect

Larsen & Toubro • Chennai District

On-site
INR 4,000,000 - 7,000,000
Machine Learning Engineer
Machine Learning Engineer

Tranzeal • Bengaluru

On-site
INR 3,500,000 - 7,500,000
Distributed Training & Inference Optimization Engineer
Distributed Training & Inference Optimization Engineer

Winzons • India

On-site
INR 3,000,000 - 5,000,000
Inference Systems Engineer
Inference Systems Engineer

Nava • Bengaluru

On-site
INR 1,700,000 - 2,500,000
AI Ml Engineer
AI Ml Engineer

Precisiontech Global It Solutions • Bengaluru, Gurugram District

Hybrid
INR 2,500,000 - 4,000,000
AI Engineer Global IT Infrastructure
AI Engineer Global IT Infrastructure

Vishay Intertechnology, Inc. • Pune District

On-site
INR 1,200,000 - 3,200,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Bhavitha Tech, CMMi Level 3 Company • Bengaluru

On-site
INR 1,200,000 - 1,800,000