We are seeking a highly skilled Lead / Principal MLOps Engineer to design, build, and operate secure, scalable,and reliable enterprise ML platform infrastructure. This role will lead the end-to-end operationalization ofmachine learning workloads—from training and experimentation through deployment, serving, monitoring, andlifecycle management.
Responsibilties
- Design, build, and maintain scalable, secure, highly available, and reliable MLOps and ML platforminfrastructure.
- Lead end-to-end ML pipelines covering training, validation, deployment, serving, monitoring, and lifecyclemanagement.
- Build and manage Kubernetes-based model deployment and serving infrastructure.
- Implement robust CI/CD pipelines for ML applications, models, and platform infrastructure.
- Manage Infrastructure as Code and automate cloud provisioning, configuration, and environmentmanagement.
- Design and optimize Kubernetes environments for high availability, scalability, disaster recovery, security,and efficient resource utilization.
- Implement monitoring and observability solutions across ML models, pipelines, applications, infrastructure,and platform services.
- Monitor and manage model performance, data drift, concept drift, reliability, and operational health.
- Optimize cloud and GPU infrastructure for performance, scalability, reliability, and cost efficiency.
- Lead troubleshooting, incident management, root cause analysis, and production issue resolution.Collaborate with Data Science, ML Engineering, Platform Engineering, and Cloud teams to improve MLworkflows.
- Establish MLOps best practices, engineering standards, architecture patterns, governance controls, anddocumentation.
- Provide technical leadership, mentoring, architecture guidance, and infrastructure/code reviews..
Required Qualification
- 8+ years of overall engineering experience, including at least 4 years of hands-on experience in MLOps, MLPlatform Engineering, or Cloud Engineering.
- Strong hands-on experience designing and operating production MLOps or ML platform environments.
- Strong understanding of ML lifecycle management: experimentation, model versioning, model registry,training, deployment, serving, monitoring, and governance.
- Hands-on experience with one or more: MLflow, Kubeflow, Apache Airflow, Argo Workflows, Weights &Biases (W&B), or Amazon SageMakerAdvanced experience with Docker and Kubernetes, including EKS, GKE, or AKS.
- Strong practical experience with Helm and KServe for Kubernetes-based application and model deployment/serving.
- Strong experience with GitHub Actions, GitLab CI, or Jenkins to automate build, test, deployment, andrelease processes.
- Strong cloud-platform experience in AWS, Azure, or GCP; multi-cloud or hybrid-cloud experience ispreferred.
- Strong experience with ML pipeline orchestration, model deployment and serving, model monitoring, datadrift and concept drift detection, autoscaling, HA, DR, incident management, and RCA.
- Experience optimizing cloud and GPU infrastructure for performance, scalability, reliability, and costefficiency.
- Strong problem-solving, technical leadership, communication, stakeholder-management, and mentoringabilities.