Technical Lead-Machine learning

Myntra

Bengaluru

Hybrid

INR 400,000 - 700,000

Full time

12 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Myntra is seeking a Technical Lead – ML Platform & MLOps to design and build the next generation infrastructure powering ML across Myntra. You will own the ML platform, enabling Data Scientists and ML Engineers to train, experiment, deploy and operate models with self-service capabilities.

The role focuses on scalable training, inference and lifecycle management across multi-node GPU/CPU clusters, with emphasis on reliability, observability and collaboration across product and engineering teams.

Qualifications

  • 6+ years of software/platform engineering experience with ML
  • Strong programming skills in Python
  • Deep hands-on experience with Kubernetes, containers and Linux
  • Experience building or operating large-scale distributed systems or platform infrastructure
  • Strong understanding of cloud infrastructure and CI/CD
  • Experience with workflow orchestration and ML lifecycle tooling

Responsibilities

  • Architect, build and evolve scalable ML Platform and MLOps capabilities for large-scale production workloads.
  • Create self-service infrastructure for Data Scientists to train, deploy and operate models.
  • Design reusable platform abstractions, SDKs, APIs and tooling to improve ML developer productivity
  • Build systems for reproducibility, lineage, versioning, governance and lifecycle management across ML workflows
  • Drive architecture and technology choices for Myntra's ML infrastructure

Skills

Python
Kubernetes
Distributed systems
CI/CD
Terraform

Tools

Airflow
Ray
MLflow
Kubernetes

Job description

Myntra is looking for a Technical Lead – ML Platform & MLOps to help build the next generation of infrastructure that powers machine learning across Myntra.

You will have the opportunity to design and build foundational ML platform capabilities used across multiple ML teams, influence platform architecture, solve large-scale infrastructure challenges, and establish engineering standards for how ML workloads are built and operated at Myntra.

We are looking for a strong hands-on engineer who enjoys building platforms, debugging complex distributed systems, and simplifying infrastructure for hundreds of ML and engineering users.

What You Will Own
Build Myntra's ML Platform
  • Architect, build and evolve scalable ML Platform and MLOps capabilities for large-scale production workloads.
  • Create self-service infrastructure that allows Data Scientists and ML Engineers to train, experiment, deploy and operate models without managing underlying infrastructure.
  • Design reusable platform abstractions, SDKs, APIs and tooling that dramatically improve ML developer productivity.
  • Build systems for reproducibility, lineage, versioning, governance and lifecycle management across ML workflows.
  • Drive architecture and technology choices for Myntra's ML infrastructure.
Distributed Training & GPU Infrastructure
  • Build infrastructure for large-scale distributed training across CPU and GPU clusters.
  • Design and operate multi-node and multi-GPU training environments.
  • Work with distributed computing frameworks such as Ray or equivalent technologies.
  • Build intelligent scheduling, resource isolation and autoscaling capabilities for ML workloads.
  • Improve utilization of expensive GPU infrastructure through scheduling, workload optimization and capacity management.
  • Design fault-tolerant training systems with checkpointing, retries and recovery mechanisms.
ML Workflow & Training Platform
  • Build a world-class training platform supporting the complete ML development lifecycle.
  • Design scalable orchestration for training, feature engineering, validation and deployment workflows.
  • Build reusable workflow components using technologies such as Airflow, Ray and Kubernetes.
  • Improve scheduling, dependency management, execution isolation and reliability for thousands of ML workloads.
  • Enable experimentation across different compute environments without exposing infrastructure complexity to users.
MLOps & Model Lifecycle
  • Build end-to-end ML lifecycle capabilities covering:
  • Build experiment tracking and model management using MLflow or equivalent technologies.
  • Enable reliable model versioning, approval, rollout and rollback.
  • Build automated model validation and production-readiness workflows.
  • Enable reproducible ML workflows across development, staging and production environments.
Model Serving & AI Infrastructure
  • Build highly scalable infrastructure for real-time, batch and asynchronous model inference.
  • Design model-serving platforms running on Kubernetes and GPU infrastructure.
  • Optimize serving systems for latency, throughput, availability and cost.
  • Explore and adopt technologies such as Ray Serve, NVIDIA Triton, vLLM, SGLang or equivalent platforms where appropriate.
  • Enable production deployment of traditional ML, deep-learning and emerging AI/LLM workloads.
Platform Reliability & Observability
  • Treat ML infrastructure as a production-grade distributed platform.
  • Define and drive SLIs, SLOs, availability and reliability standards for ML platform services.
  • Build deep observability across infrastructure, pipelines, training workloads and inference systems.
  • Troubleshoot challenging production issues spanning Kubernetes, GPU workloads, distributed systems, networking, storage and ML pipelines.
  • Drive root-cause analysis and systematically eliminate recurring operational issues.
  • Design for high availability, fault tolerance and graceful recovery.
  • Design scalable compute, networking and storage infrastructure for ML workloads.
  • Build and operate ML systems on Kubernetes and cloud platforms.
  • Automate infrastructure using Terraform or equivalent Infrastructure-as-Code technologies.
  • Build secure, isolated and reproducible runtime environments.
  • Drive infrastructure efficiency through autoscaling, workload placement and cost optimization.

A major part of this role is making complex ML infrastructure simple for users.

You Will:
  • Build developer-facing platforms, SDKs, APIs and abstractions.
  • Reduce the time required to move an ML experiment into production.
  • Eliminate repetitive infrastructure work for Data Scientists and ML Engineers.
  • Build standardized templates and paved roads for ML development.
  • Improve debugging, discoverability and observability of ML workloads.
  • Enable teams to focus on models and business problems rather than infrastructure.
Technical Leadership

At E3, we expect you to go beyond implementing individual components.

You Will:
  • Own architecture and technical direction for major ML Platform initiatives.
  • Lead complex system-design discussions and technical reviews.
  • Convert ambiguous problems into scalable platform solutions.
  • Drive engineering excellence across reliability, scalability, performance and maintainability.
  • Mentor engineers and raise the technical bar of the team.
  • Influence architecture across ML, Data, Platform, SRE and Infrastructure teams.
  • Evaluate emerging technologies and make pragmatic build-vs-buy decisions.
  • Take critical systems from concept through architecture, implementation and production adoption.
What We Are Looking For
Must Have
  • 6+ years of strong hands-on software/platform engineering experience.
  • Strong programming skills in Python.
  • Deep hands-on experience with Kubernetes, containers and Linux.
  • Experience building or operating large-scale distributed systems or platform infrastructure.
  • Strong understanding of cloud infrastructure including compute, storage and networking.
  • Experience building production-grade CI/CD and automation platforms.
  • Experience with workflow orchestration such as Apache Airflow or equivalent systems.
  • Experience with ML lifecycle tooling such as MLflow or equivalent platforms.
  • Strong understanding of ML training and deployment workflows.
  • Experience with Infrastructure as Code, preferably Terraform.
  • Strong debugging and production troubleshooting skills.
  • Experience building systems with monitoring, logging, metrics and alerting.
  • Strong fundamentals in system design, reliability and distributed computing.
Strong Differentiators

We would especially love to meet you if you have worked on:

  • Ray or other distributed computing frameworks
  • Distributed or multi-node ML training
  • GPU and multi-GPU infrastructure
  • ML training platforms used by multiple teams
  • Model serving and inference infrastructure
  • GPU scheduling and utilization optimization
  • Large-scale workflow orchestration
  • Infrastructure cost and performance optimization
  • Feature platforms or feature stores
Good to Have
  • Databricks, SageMaker, Vertex AI or similar ML platforms
  • Model monitoring, data drift and automated retraining
  • NVIDIA Triton, Ray Serve, vLLM or SGLang
  • LLM training/inference and LLMOps
  • Vector databases and retrieval infrastructure
  • Model governance and lineage
  • OpenTelemetry, Grafana or similar observability ecosystems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Lead – ML Platform & MLOps
Technical Lead – ML Platform & MLOps

Myntra • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Machine Learning Platform Engineer
Senior Machine Learning Platform Engineer

Amgen SA • Hyderabad

On-site
INR 4,000,000 - 7,000,000
MLOps Manager
MLOps Manager

Anblicks • Hyderabad

On-site
INR 2,000,000 - 3,000,000
ML platform Engineer ( AI control plane)
ML platform Engineer ( AI control plane)

Talentiser • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior Machine Learning Platform Engineer
Senior Machine Learning Platform Engineer

Amgen • Hyderabad

On-site
INR 3,500,000 - 6,500,000
Senior Software Engineer ( AI/ML Developer )
Senior Software Engineer ( AI/ML Developer )

Lyric • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Technical Architect - ML
Technical Architect - ML

Prodapt Solutions Private Limited • Chennai District

On-site
INR 5,000,000 - 7,500,000
Technical Architect - ML
Technical Architect - ML

Prodapt • Chennai District

On-site
INR 4,000,000 - 6,000,000
AI-ML Engineer
AI-ML Engineer

KanthamAi • Mumbai

On-site
INR 2,000,000 - 3,000,000
Machine Learning Engineer - 2
Machine Learning Engineer - 2

SatSure Analytics India • India

On-site
INR 1,500,000 - 4,000,000