Lead MLOps

Rakuten Symphony

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Responsibilities include designing feature stores, model monitoring, and scalable data pipelines while mentoring engineers and ensuring platform adoption across enterprise customers. The role emphasizes cloud-native practices and collaboration with security and product teams.

Qualifications

  • 7+ years in software engineering, data engineering, or MLOps with 5+ years in lead/architect role.
  • Expert-level Kubernetes and ecosystem knowledge (operators, CRDs, Helm, networking, storage).
  • Experience with ML platform tools such as MLflow, Kubeflow, Airflow, SageMaker, Vertex AI, or Azure ML.
  • Experience with Kubeflow in production (Pipelines, Notebooks, Training Operators, KServe).
  • Designs feature stores (e.g., Feast) for online/offline serving; data processing with Spark on Kubernetes.

Responsibilities

  • Lead architectural vision and evolution of the platform aligning with business goals, security, and scalability.
  • Oversee integration of MLOps tools into a cohesive enterprise solution; guide DevOps/Platform and MLOps teams.
  • Provide guidance on Kubernetes-native MLOps practices and model serving strategies (KServe).
  • Architect enterprise security features: SSO, secrets management, RBAC, and policy enforcement for data and resources.
  • Mentor and upskill engineering teams in MLOps best practices and cloud-native development.
  • Collaborate with multiple stakeholders to gather requirements and drive platform adoption.

Skills

Leadership
Strategic thinking
Communication
Mentoring
Stakeholder management

Tools

Kubernetes
Kubeflow
MLflow
Spark
KServe
Keycloak
HashiCorp Vault
OPA Gatekeeper
Prometheus
Grafana
ELK/OpenSearch
Fluentd
Jaeger
Python
TensorFlow
PyTorch
Scikit-learn
MinIO
S3
GCS
Iceberg
Delta Lake

Job description

Job Title: MLOps Lead - Enterprise AI Platform

Responsibilities
Strategic Platform Architecture
  • Lead the architectural vision, design, and continuous evolution of the Platform, ensuring alignment with business objectives, security standards, and scalability requirements.
  • Drive the adoption and integration of open-source MLOps tools (Kubeflow, MLflow, Feast, KServe, Alibi-Detect, Evidently AI, Spark, etc.) into a cohesive, production-ready enterprise solution.
  • Define platform standards, best practices, and architectural patterns for MLOps development and operations.
Technical Leadership & Implementation Oversight
  • Act as the primary technical authority and lead for the MLOps initiative, guiding both DevOps/Platform and MLOps/Data Science teams through the phased development plan.
  • Oversee the implementation of core platform components, ensuring robust integration, performance, and adherence to architectural blueprints.
  • Provide expert guidance on Kubernetes-native MLOps practices, distributed computing for ML (Spark, Kubeflow Training Operators), and model serving strategies (KServe).
Enterprise Security, Governance & Multi-Tenancy
  • Architect and oversee the implementation of enterprise-grade security features including SSO (Keycloak), secrets management (HashiCorp Vault), and fine-grained access control (Kubernetes RBAC, OPA Gatekeeper) for data and platform resources.
  • Design and enforce multi-tenancy models that provide strong isolation, resource governance, and secure data access for internal teams and external customers.
  • Ensure the platform meets stringent compliance requirements through comprehensive audit logging, tracing (Fluentd, ELK/OpenSearch, Prometheus/Grafana), and data lineage considerations.
ML Lifecycle & Data Management Expertise
  • Architect and integrate a robust Feature Store (tool like Feast) for consistent feature engineering, management, and serving across training and inference.
  • Lead the integration of MLflow for experiment tracking, model versioning, and a centralized model registry.
  • Design and implement comprehensive model monitoring solutions, including data drift and model quality detection (Alibi-Detect/Evidently AI), with integrated alerting.
Developer Experience & Customization
  • Champion the developer experience for data scientists, ensuring ease of use, self-service capabilities, and efficient workflows (e.g., automated namespace provisioning, notebook environment management).
  • Provide architectural guidance for building a custom, branded UI layer on top of the open source components, enhancing usability and aligning with product offerings.
  • Collaborate extensively with Data Science, DevOps, Security, Product Management, and Business stakeholders to gather requirements, communicate technical vision, and drive platform adoption.
  • Mentor and upskill engineering teams in MLOps best practices, cloud-native development, and advanced ML techniques.
Required Skills & Expertise
  • 7+ years of progressive experience in software engineering, data engineering, or MLOps, with at least 5 years in a lead or architect role focused on building and managing production of large-scale ML platforms.
  • Expert-level proficiency with Kubernetes and its ecosystem (operators, CRDs, Helm, networking, storage).
  • Experience in building/managing ML platform tools such as MLflow , Kubeflow, Airflow, SageMaker, Vertex AI, or Azure Machine Learning.
  • Deep hands-on experience with Kubeflow (Pipelines, Notebooks, Training Operators, KServe) in production environments.
  • Extensive experience with MLflow for experiment tracking, model registry, and model lifecycle management.
  • Proven expertise in designing and implementing Feature Stores (e.g., Feast) for both online and offline serving.
  • Strong background in distributed data processing technologies like Apache Spark/PySpark, especially on Kubernetes.
  • Architectural experience with enterprise security solutions including SSO (Keycloak, OAuth/OIDC), secrets management (HashiCorp Vault), and policy enforcement (Kubernetes RBAC, OPA Gatekeeper).
  • Demonstrated ability to implement comprehensive monitoring and observability stacks (Prometheus, Grafana, ELK/OpenSearch, Fluentd, Jaeger) for platform health and ML model performance/drift (Alibi-Detect, Evidently AI).
  • Proficiency in Python and experience with major ML/Deep Learning frameworks (TensorFlow, PyTorch, Scikit-learn).
  • Experience with cloud-native storage solutions (e.g., MinIO, S3, GCS) and open table formats (Iceberg, Delta Lake).
  • Excellent communication, leadership, and interpersonal skills with the ability to influence technical direction and drive complex initiatives across multiple teams.
RAKUTEN SHUGI PRINCIPLES

Our worldwide practices describe specific behaviours that make Rakuten unique and united across the world. We expect Rakuten employees to model these 5 Shugi Principles of Success.

  • Always improve, always advance. Only be satisfied with complete success - Kaizen.
  • Be passionately professional. Take an uncompromising approach to your work and be determined to be the best.
  • Hypothesize - Practice - Validate - Shikumika. Use the Rakuten Cycle to success in unknown territory.
  • Maximize Customer Satisfaction. The greatest satisfaction for workers in a service industry is to see their customers smile.
  • Speed!! Speed!! Speed!! Always be conscious of time. Take charge, set clear goals, and engage your team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Engineer – AI/ML
Principal Engineer – AI/ML

Rakuten Symphony • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Technical Lead, AI/ML
Technical Lead, AI/ML

Rakuten Symphony • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Principal Platform Engineer
Principal Platform Engineer

Rakuten India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Architect - DevOps and ML Op's
Senior Architect - DevOps and ML Op's

Metaplore Solutions Pvt Ltd • Bengaluru

On-site
INR 3,500,000 - 6,000,000
MLOps Architect
MLOps Architect

Anblicks • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Cutting-edge technologies
Mentorship opportunities
Impactful initiatives in AI and Data
Staff MLOps Engineer
Staff MLOps Engineer

GoTo Meeting • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Sodexo Meal Coupon
Internet Reimbursement
Mobile Reimbursement
+3
Sr. Manager, Enterprise Systems
Sr. Manager, Enterprise Systems

Skyworks Solutions, Inc. • Bengaluru

On-site
Competitive salary
Career growth opportunities
Referral bonus program of Rs200,000
Senior MLOps Engineer
Senior MLOps Engineer

Vitric Business Solutions • Pune District, Mumbai, Indore District

On-site
INR 180,000 - 300,000
Chief Data and Artificial Intelligence Officer
Chief Data and Artificial Intelligence Officer

Rakuten Symphony • India

On-site
INR 6,000,000 - 12,000,000
Leadership development
MLOps Manager
MLOps Manager

Anblicks • Hyderabad

On-site
INR 2,000,000 - 3,000,000