Get more replies from employers
Send a job-specific resume in minutes.
Winzons in India is assembling an advanced AI infrastructure team focused on scalable machine learning systems, optimizing both training and inference for large models.
You will work on GPU-based workloads, improve performance of distributed training with PyTorch and DeepSpeed, and pursue low latency through quantization, caching, and model parallelism. The role spans cross-team collaboration and cutting-edge compute platforms.
Join a highly advanced AI infrastructure team focused on building and optimizing large-scale machine learning systems. This environment leverages cutting-edge technologies to enable high-performance experimentation, scalable model deployment, and efficient processing of large datasets.The team operates globally, bringing together engineers and researchers to push the boundaries of deep learning, distributed systems, and next-generation compute platforms.
This position is centered on maximizing the efficiency and scalability of GPU-based machine learning workloads, particularly for large language models (LLMs) and generative AI systems.You will work on improving both training performance and inference efficiency, ensuring optimal utilization of hardware resources, reduced latency, and faster model iteration cycles. The role requires hands-on expertise in deep learning frameworks, distributed systems, and performance optimization.