About The Role
The MLOps Engineer owns the infrastructure and automation required to move machine learning and GenAI systems from experimentation into reliable production environments. The role builds repeatable pipelines for training, evaluation, deployment, monitoring, and rollback across cloud-based compute and data platforms.
You will work with ML engineers, data scientists, and platform teams to improve model release velocity without compromising security, reproducibility, or operational reliability. The work spans Kubernetes, CI/CD, model registries, observability, and the serving systems that support latency-sensitive AI products.
Key Responsibilities
- Build and maintain ML pipelines for data validation, feature processing, model training, evaluation, registration, and promotion using tools such as Kubeflow, Airflow, MLflow, or equivalent
- Automate model and service deployment with Docker, Kubernetes, Helm, and Terraform across AWS, Azure, or GCP environments
- Develop CI/CD workflows for Python services, model artifacts, infrastructure changes, and container images using GitHub Actions, GitLab CI, or Jenkins
- Implement production monitoring for service health, latency, throughput, resource utilization, data drift, model performance, and model quality regressions
- Operate scalable model-serving infrastructure using platforms such as KServe, Seldon, NVIDIA Triton, Ray Serve, or managed cloud endpoints
- Establish reproducibility and governance practices for datasets, features, model versions, experiments, secrets, and deployment approvals
- Troubleshoot production incidents, improve system reliability, and document runbooks, architecture decisions, and operational standards
What We Are Looking For
- 3–8 years of experience in MLOps, platform engineering, DevOps, or software engineering supporting machine learning systems in production
- Strong Python and Linux skills, with experience building APIs, automation tools, and production services
- Hands-on experience with Docker, Kubernetes, infrastructure as code, and at least one major cloud platform: AWS, Azure, or GCP
- Practical knowledge of ML lifecycle tooling such as MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, Airflow, or comparable platforms
- Experience designing CI/CD pipelines and observability using tools such as Prometheus, Grafana, OpenTelemetry, ELK, or cloud-native monitoring services
- Bachelor’s degree in computer science, engineering, mathematics, or a related technical field; equivalent professional experience will be considered
- Bonus: Experience with LLM serving, GPU scheduling, Ray, NVIDIA Triton, feature stores, GitOps, service meshes, model governance, or regulated production environments