Distributed Training & Inference Optimization Engineer

Winzons

India

On-site

INR 3,000,000 - 5,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Winzons in India is assembling an advanced AI infrastructure team focused on scalable machine learning systems, optimizing both training and inference for large models.

You will work on GPU-based workloads, improve performance of distributed training with PyTorch and DeepSpeed, and pursue low latency through quantization, caching, and model parallelism. The role spans cross-team collaboration and cutting-edge compute platforms.

Qualifications

  • 3+ years optimizing GPU-based ML workloads.
  • Strong experience with PyTorch, DeepSpeed, or equivalents.
  • Experience in distributed training for large-scale models.
  • Knowledge of inference optimization: quantization, pruning, caching.
  • Degree in CS/Engineering or related technical field.

Responsibilities

  • Enhance performance of distributed training frameworks such as PyTorch, DeepSpeed, or similar systems.
  • Identify and resolve bottlenecks in large-scale training pipelines (e.g., memory usage, communication overhead, GPU utilization).
  • Optimize inference systems using techniques like quantization, caching, and batching to achieve low latency and high throughput.
  • Collaborate with infrastructure and platform teams to improve resource orchestration, scheduling, and system reliability.
  • Design benchmarking tools and metrics to measure training efficiency, system throughput, and latency performance.
  • Apply advanced optimization techniques (e.g., mixture-of-experts, speculative decoding, model parallelism) to improve large model performance.

Skills

GPU optimization
PyTorch
DeepSpeed
Distributed training
Inference optimization
CUDA

Education

Bachelor's degree in Computer Science or Engineering

Tools

CUDA profiling tools
Containerization (Docker)
Model optimization techniques (e.g., FlashAttention, LoRA)

Job description

Overview

Join a highly advanced AI infrastructure team focused on building and optimizing large-scale machine learning systems. This environment leverages cutting-edge technologies to enable high-performance experimentation, scalable model deployment, and efficient processing of large datasets.The team operates globally, bringing together engineers and researchers to push the boundaries of deep learning, distributed systems, and next-generation compute platforms.


About the Role

This position is centered on maximizing the efficiency and scalability of GPU-based machine learning workloads, particularly for large language models (LLMs) and generative AI systems.You will work on improving both training performance and inference efficiency, ensuring optimal utilization of hardware resources, reduced latency, and faster model iteration cycles. The role requires hands-on expertise in deep learning frameworks, distributed systems, and performance optimization.


Key Responsibilities


  • Enhance performance of distributed training frameworks such as PyTorch, DeepSpeed, or similar systems

  • Identify and resolve bottlenecks in large-scale training pipelines (e.g., memory usage, communication overhead, GPU utilization)

  • Optimize inference systems using techniques like quantization, caching, and batching to achieve low latency and high throughput

  • Collaborate with infrastructure and platform teams to improve resource orchestration, scheduling, and system reliability

  • Design benchmarking tools and metrics to measure training efficiency, system throughput, and latency performance

  • Apply advanced optimization techniques (e.g., mixture-of-experts, speculative decoding, model parallelism) to improve large model performance

  • Continuously evaluate new approaches to hardware acceleration and model execution efficiency


Required Qualifications


  • 3+ years of hands-on experience optimizing GPU-based machine learning workloads

  • Strong expertise in deep learning frameworks such as PyTorch, DeepSpeed, or equivalent

  • Experience with distributed training techniques for large-scale models

  • Solid understanding of inference optimization strategies (e.g., quantization, pruning, caching, batching)

  • Degree in Computer Science, Engineering, or a related technical field


Preferred Qualifications


  • Experience with CUDA programming and GPU performance profiling tools

  • Familiarity with distributed systems communication libraries and optimization techniques

  • Knowledge of model optimization methods such as FlashAttention, LoRA, or similar techniques

  • Experience working with containerized or orchestrated environments for ML workloads
  • Contributions to open-source machine learning or infrastructure projects

  • Hands-on experience with modern inference serving frameworks

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Inference Engineer
Senior Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 2,600,000 - 4,800,000
Inference Engineer
Inference Engineer

Binaire Private Limited • New Delhi

On-site
INR 800,000 - 1,200,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

LeadSoc Technologies Pvt Ltd • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Edge AI deployment
Generative AI systems
Distributed inference pipelines
Principal Research Engineer, Applied AI
Principal Research Engineer, Applied AI

Mulya Technologies • India

On-site
INR 1,800,000 - 3,000,000
Principal Research Engineer, Applied AI
Principal Research Engineer, Applied AI

EnCharge AI • India

On-site
INR 3,000,000 - 6,000,000
Lead MLOps Engineer
Lead MLOps Engineer

Cloudkeeper • Garhi

On-site
INR 3,800,000 - 6,400,000
AI Infrastructure and Platform Architect
AI Infrastructure and Platform Architect

Ignatiuz Inc. • Indore District

On-site
INR 3,500,000 - 6,000,000
Senior AI Software Performance Engineer
Senior AI Software Performance Engineer

BigStep Technologies • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Senior System Software Engineer - LocalAI
Senior System Software Engineer - LocalAI

NVIDIA Gruppe • Pune District

On-site
INR 300,000 - 550,000
Senior System Software Engineer - LocalAI
Senior System Software Engineer - LocalAI

NVIDIA • Pune District

On-site
INR 4,000,000 - 7,000,000