An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Tata Consultancy Services in Kolkata is seeking a seasoned MLOps Engineer to design and manage automated pipelines for deploying machine learning models into production. You will ensure seamless collaboration between data science, data engineering and application teams, implement versioning/rollback, and build scalable cloud-native infrastructure.
Role emphasizes collaboration with DevOps, setting up CI/CD for ML and maintaining monitoring, security, and performance optimization across models
- Design, develop, and manage automated pipelines for deploying machine learning models into production.
- Ensure smooth integration between model development, data, and application teams.
- Implement model versioning and rollback strategies to facilitate easy model updates and troubleshooting.
- Build and maintain scalable infrastructure using tools like Kubernetes, Docker, and cloud platforms (AWS, Azure, GCP).
- Automate the deployment process and manage model serving environments.
- Design and optimize cloud-native solutions to ensure scalability and performance under heavy workloads.
- Continuously monitor the performance and health of deployed models in production environments.
- Implement real-time logging, alerting, and monitoring systems to ensure models effectiveness over time.
- Detect, troubleshoot, and resolve issues such as model drift, degradation, and inefficiencies.
- Work closely with data scientists to ensure that models are production-ready and meet system requirements.
- Collaborate with DevOps teams to integrate MLOps tools and practices into the CI/CD pipeline.
- Optimize model performance by coordinating with various teams to manage the lifecycle of machine learning models.
- Automate and manage model retraining processes based on incoming new data or changing business needs.
- Create frameworks for evaluating and improving model accuracy, efficiency, and robustness.
- Ensure the security of machine learning systems, including data protection, model access control, and sensitive data handling.
- Ensure compliance with relevant regulatory requirements related to data privacy and security.
- Work on optimizing models and system performance for faster inference and low-latency predictions.
- Implement techniques like quantization, pruning, and model distillation to optimize the models runtime efficiency.
- Maintain comprehensive documentation for model deployment pipelines, monitoring setups, and operational procedures.
- Provide regular reports on system performance, model health, and operational metrics to stakeholders.
- Stay up-to-date with emerging MLOps technologies and best practices.
- Research and implement new tools and frameworks to improve operational efficiency
Bachelors or Masters degree in Computer Science, Engineering, Mathematics, or a related field.
- Strong experience with cloud platforms (AWS, Google Cloud, Azure).
- Familiarity with Docker, Kubernetes, and containerization technologies
- Proficiency in programming languages such as Python, Java, or Go
- Experience with CI/CD tools (Jenkins, GitLab CI, etc.) and automation frameworks.
- Familiarity with version control systems (e.g., Git).
- Working knowledge of machine learning frameworks (TensorFlow, PyTorch, Scikit-learn, etc.).
- Experience with model serving tools like TensorFlow Serving, TorchServe, or MLFlow.
- Strong knowledge of data pipelines and ETL processes.
- Experience with Big Data technologies (Spark, Hadoop, Kafka) is a plus.
- Expertise in data preprocessing and **feature engineering.
- Proficiency with monitoring tools like Prometheus, Grafana, or Datadog.
- Experience with logging frameworks such as ELK stack (Elasticsearch, Logstash, Kibana) or Splunk.
- Strong understanding of **software engineering principles** and best practices.
- Experience in designing highly available, fault-tolerant, and scalable distributed systems.