Aligned Automation Services Pvt Ltd | Full time
At Aligned Automation, we livebyour"Better Together"philosophy to build a betterworld. As a strategic service provider to Fortune 500 companies, we helpdigitize enterprise operations and drive impactful business strategies. Our purposegoes beyond projects—we strive to deliver meaningful, sustainable change thatshapes a more optimistic and equitable future.
Our culture is deeply rooted inour4Cs—Care,Courage,Curiosity,andCollaboration—ensuring that each employee is empowered to grow,innovate, and thrive in an inclusive workplace.
About the Role
Key Responsibilities
- Design, build, and maintain scalable MLOps platforms for training, deploying, monitoring, and managing machine learning models.
- Develop and optimize CI/CD pipelines for ML applications using GitLab CI, Jenkins etc.
- Containerize applications using Docker and orchestrate workloads using Kubernetes.
- Build and maintain ML pipelines using Kubeflow, MLflow, Airflow, Argo Workflows, or similar orchestration tools.
- Implement model versioning, experiment tracking, model registry, and reproducible ML workflows.
- Deploy ML models as REST/gRPC APIs using FastAPI, Flask, or similar frameworks.
- Configure monitoring, logging, and alerting using Prometheus, Grafana, Dynatrace, or Splunk.
- Optimize model serving performance, scalability, and resource utilization.
- Work closely with Data Scientists, Data Engineers, Platform Engineers, and Software Developers to productionize ML solutions.
- Ensure platform security, governance, and compliance following DevSecOps best practices.
- Troubleshoot production issues and perform root cause analysis for ML workloads.
Required Technical Skills
- Strong programming experience in Python.
- Hands‑on experience with Docker, Kubernetes, Helm, and containerized deployments.
- Strong knowledge of Git, branching strategies, merge requests, and CI/CD.
- Experience with MLflow, Kubeflow, Airflow, or similar MLOps tools.
- Experience with Spark/PySpark and distributed data processing.
- Knowledge of REST APIs, FastAPI, Flask, and microservices architecture.
- Experience with Linux, Bash scripting, and automation.
- Familiarity with model monitoring, drift detection, feature stores, and model governance.
- Experience working with SQL and NoSQL databases.
- Strong understanding of software engineering principles, design patterns, testing, and code quality.
Preferred Skills
- Experience with Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), vector databases, and AI orchestration frameworks such as LangChain or langfuse.
- Experience with NVIDIA GPU workloads and model optimization.
- Knowledge of Kafka, RabbitMQ, or event‑driven architectures.
- Exposure to feature stores such as Feast.
- Experience with OpenShift, Red Hat ecosystem, or enterprise Kubernetes platforms.
- Knowledge of security scanning tools such as Trivy, SonarQube, and vulnerability management.