We are seeking an experienced and highly motivated MLOps Engineer to join the Data and AI team. In this role, you will bridge the gap between machine learning development and scalable production systems by building, automating, and managing end-to-end ML pipelines. The ideal candidate will have strong expertise in cloud platforms, CI/CD automation, infrastructure-as-code, and productionizing machine learning models in enterprise environments.
Key Responsibilities
- Design, build, and maintain scalable ML infrastructure and CI/CD pipelines for training, testing, deploying, and monitoring machine learning models
- Automate model versioning, deployment, rollback strategies, and environment management across staging and production
- Collaborate closely with Data Scientists and Machine Learning Engineers to productionize ML models and optimize deployment workflows
- Apply Infrastructure-as-Code (IaC) practices to provision and manage cloud-based ML infrastructure
- Implement monitoring, logging, and alerting solutions for ML systems, including model drift and data anomaly detection
- Optimize the performance, scalability, and reliability of model training and inference systems
- Ensure ML operations adhere to organizational security, compliance, and reliability standards
- Maintain comprehensive documentation for systems, workflows, processes, and operational procedures
- Support continuous improvement initiatives related to MLOps, DevOps, and AI/ML operational practices
Required Qualifications
- Bachelor’s Degree in Computer Science, Engineering, or a related field
- 3+ years of experience in MLOps, DevOps, or Machine Learning Engineering
- Hands‑on experience with Azure DevOps and Azure Machine Learning (AzureML)
- Proficiency with cloud platforms such as AWS, Azure, or GCP
- Experience with containerization and orchestration technologies including Docker and Kubernetes
- Strong programming skills in Python, Bash, and PowerShell
- Experience working with REST APIs
- Experience with Infrastructure-as-Code tools such as Terraform or ARM templates
- Familiarity with CI/CD tools including Jenkins, GitHub Actions, or Azure DevOps Pipelines
- Hands‑on experience with machine learning frameworks such as TensorFlow, PyTorch, and Scikit‑learn
- Familiarity with ML tools such as MLflow, TFX, DVC, or Kubeflow
- Experience with workflow orchestration tools such as Apache Airflow or Prefect
- Strong troubleshooting, analytical, and problem‑solving skills
- Excellent verbal and written communication skills
Preferred Qualifications
- Experience with monitoring and logging tools such as Azure Monitor, Prometheus, or Grafana
- Experience working in Agile development environments
- Experience collaborating with Data Science and Machine Learning teams
- Strong collaboration skills with experience working in cross‑functional technical teams
- Experience optimizing scalable AI/ML operations and infrastructure
Certifications
- Certified Kubernetes Administrator (CKA) or equivalent certification