AI Platform Engineer (Inference)

Amazon

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon Bengaluru seeks an experienced backend and distributed-systems engineer to scale AI inference platforms. You will use Python or Go to design reliable services and help deploy GPU-accelerated workloads.

Ideal candidates will know inference frameworks (vLLM, Triton, TensorRT-LLM, ONNX Runtime, PyTorch), distributed training concepts, and cloud-native tools like Kubernetes and container orchestration. Join a team driving high-performance ML at scale.

Qualifications

  • 9+ years of backend development / distributed systems experience with Python or Go.
  • Hands-on with inference frameworks: vLLM, Triton, TensorRT-LLM, ONNX Runtime, PyTorch.
  • Understanding of distributed training: data parallelism, tensor parallelism, pipeline parallelism, ZeRO.
  • Knowledge of Kubernetes and cloud-native ecosystem; containerization and service orchestration.
  • Proficiency in GPU cluster management fundamentals (CUDA, cuDNN, device plugins).
  • Strong distributed systems design: load balancing, HA, resource scheduling.

Skills

Backend development
Distributed systems
Python
Go
Inference frameworks
Kubernetes
Cloud-native
GPU cluster management
Distributed training
Load balancing

Tools

vLLM
Triton
TensorRT-LLM
ONNX Runtime
PyTorch
CUDA
cuDNN
Docker

Job description

Requirements:

  • 9+ years of backend development / distributed systems; Python or Go.
  • Hands-on experience with one or more inference frameworks: vLLM, Triton, TensorRT-LLM, ONNX Runtime, PyTorch.
  • Understanding of distributed training: data parallelism, tensor parallelism, pipeline parallelism, ZeRO.
  • Knowledge of Kubernetes and cloud-native ecosystem; containerization and service orchestration.
  • Proficiency in GPU cluster management fundamentals (CUDA, cuDNN, device plugins).
  • Strong distributed systems design: load balancing, HA, resource scheduling.

Nice-to-Have:

  • LLM serving at scale (vLLM, SGLang, TensorRT-LLM) in production.
  • Quantisation techniques: INT8/FP8 sparsity, speculative decoding.
  • Large-scale GPU training cluster operation (Slurm, Volcano, KubeRay).
  • Training fault recovery, checkpoint management, elastic training.
  • RDMA/NCCL/InfiniBand high-performance networking.
  • Active open-source contributions (vLLM, SGLang, Megatron-LM, DeepSpeed, Ray).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Systems Engineer
Inference Systems Engineer

Nava • Bengaluru

On-site
INR 1,700,000 - 2,500,000
Inference Engineer
Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 800,000 - 1,200,000
Senior Inference Engineer
Senior Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 2,600,000 - 4,800,000
AI Inference Engineer – LLM
AI Inference Engineer – LLM

GyanSys Inc. • Bengaluru

On-site
INR 1,000,000 - 1,600,000
Distributed Training & Inference Optimization Engineer
Distributed Training & Inference Optimization Engineer

Winzons • India

On-site
INR 3,000,000 - 5,000,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Bengaluru

Hybrid
INR 5,500,000 - 9,000,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Bengaluru

On-site
INR 3,500,000 - 7,500,000
Machine Learning Engineer
Machine Learning Engineer

Tranzeal • Bengaluru

On-site
INR 3,500,000 - 7,500,000
Software Engineer II - Serverless Inference
Software Engineer II - Serverless Inference

Jobtailor • Bengaluru

On-site
INR 2,400,000 - 4,200,000
Lead MLOps Engineer
Lead MLOps Engineer

Cloudkeeper • Garhi

On-site
INR 3,800,000 - 6,400,000