Job Summary
Build, automate, and scale intelligent systems that move seamlessly from experimentation to reliable production. In this role, youll work at the intersection of DevOps and MLOpshelping teams ship ML-powered features faster, safer, and with measurable impact. Youll partner closely with data scientists, engineers, and platform teams to create repeatable pipelines, production-grade deployments, and strong observability across environments. If you enjoy solving real-world reliability challenges, improving developer experience through automation, and enabling ML models to perform consistently in production, this is a great opportunity to grow your ownership and technical depth while contributing to a collaborative, high-learning culture.
Responsibilities Platform Automation
- Design and maintain CI/CD workflows to automate build, test, release, and deployment processes for ML and supporting services.
- Implement infrastructure automation and configuration management to ensure consistent environments across dev, staging, and production.
- Improve system reliability through monitoring, alerting, incident response practices, and post-incident improvements.
MLOps Model Delivery
- Build and manage ML pipelines for training, validation, packaging, and deployment with reproducibility and traceability.
- Enable model versioning, artifact management, and controlled rollouts (e.g., canary/blue-green) for ML services.
- Establish model performance monitoring, drift detection signals, and feedback loops for continuous improvement.
Collaboration Engineering Excellence
- Work with data science teams to productionize Python ML code with robust testing, packaging, and runtime optimization.
- Define operational standards (logging, metrics, SLOs) and contribute to documentation and runbooks.
- Participate in code reviews and propose improvements to security, scalability, and cost efficiency.
Minimum Qualifications
- BTech / MTech / MCA / MSc (or equivalent practical experience).
- 23 years of hands-on experience in DevOps and/or MLOps-focused engineering roles.
- Working experience with CI/CD concepts and automation for deployments and releases.
- Practical experience supporting Python-based ML workloads (packaging, environments, dependency management, runtime troubleshooting).
- Strong understanding of Linux fundamentals, networking basics, and system troubleshooting.
Technical Requirements Skills
Good to have skills
- Docker
- Kubernetes
- Terraform
- MLflow
- Airflow
Preferred Qualifications
- Experience productionizing ML workflows end-to-end (training pipelines, model registry/artifacts, deployment, monitoring).
- Exposure to containerization and orchestration for scalable ML services (e.g., Docker, Kubernetes).
- Familiarity with Infrastructure as Code and configuration tools (e.g., Terraform, Ansible).
- Experience with ML lifecycle tooling (e.g., MLflow, Kubeflow) and workflow orchestration (e.g., Airflow).
- Hands-on exposure to LLM-enabled applications, including deployment patterns, inference optimization, and evaluation/monitoring approaches.
- Strong communication skills to align platform practices across engineering and data science stakeholders.
Educational Requirement
- MCA, MSc, MTech, Bachelor of Engineering, BTech
Preferred Skills
- Technology->Devops->Ansible
- Technology->AI-AI Engineering->MLOps
- Technology->OpenSystem->Python - OpenSystem
- Technology->AI-Data science->Machine Learning
Service Line
Data Analytics Unit