Machine Learning Ops. Engineer

Navlakha Management Services

Mumbai

On-site

INR 2,400,000 - 4,200,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Navlakha Management Services is seeking an ML/AI Ops Engineer to lead end-to-end deployment of ML models and AI agents into production. You will design CI/CD pipelines for GenAI workflows and optimize GPU-based workloads in cloud or on-prem environments.

You will ensure security, RBAC, data encryption, and regulatory compliance, while integrating monitoring and observability across distributed systems.

Qualifications

  • Strong programming skills in Python and scripting languages.
  • Hands-on experience with CI/CD tools (Jenkins, GitHub Actions, GitLab CI, Azure DevOps, etc.).
  • Experience with Docker, Kubernetes, Helm.
  • Knowledge of ML lifecycle tools (MLflow, Kubeflow, Airflow, etc.).
  • Experience with GPU optimization and distributed training frameworks.
  • Familiarity with cloud platforms (AWS/Azure/GCP) and hybrid infrastructure.
  • Experience managing vector databases.
  • Knowledge of monitoring tools (Prometheus, Grafana, ELK, Datadog, etc.).
  • Understanding of data security, encryption, compliance frameworks (GxP, HIPAA preferred).

Responsibilities

  • Own end-to-end deployment of ML models and AI agents into production environments.
  • Manage model versioning, rollback strategies, and lifecycle management.
  • Ensure high availability, scalability, and reliability of deployed systems.
  • Design and implement CI/CD pipelines tailored for ML and GenAI workflows.
  • Automate model training, testing, validation, and deployment processes.
  • Integrate model testing frameworks, performance checks, and compliance gates into pipelines.
  • Enable seamless integration between development, staging, and production environments.
  • Manage and optimize GPU-based workloads for model training and inference.
  • Monitor and improve compute utilization, cost efficiency, and latency.
  • Administer and maintain cloud (AWS/Azure/GCP) and/or on-prem infrastructure.
  • Support containerized deployments using Docker and Kubernetes.
  • Implement infrastructure and model performance monitoring systems.
  • Track system health, latency, throughput, resource utilization, and failure rates.
  • Establish alerting, logging, and incident response processes.
  • Continuously improve system performance and reliability through proactive monitoring.
  • Deploy and manage vector databases (e.g., FAISS, Pinecone, Weaviate, Chroma).
  • Optimize indexing, embedding pipelines, and retrieval performance for GenAI apps.
  • Ensure high availability and backup strategies for AI data systems.
  • Implement RBAC and secure authentication mechanisms.
  • Ensure infra and AI systems comply with pharma regulatory standards and internal governance policies.
  • Manage sensitive data securely, including encryption (at rest and in transit).
  • Support audit readiness and documentation for compliance reviews.

Skills

Python
CI/CD tools
Docker
Kubernetes
ML lifecycle tools
GPU optimization
Cloud platforms
Vector databases
Monitoring tools
Security/compliance

Tools

Jenkins
GitHub Actions
GitLab CI
Azure DevOps
Helm
MLflow
Kubeflow
Airflow
Prometheus
Grafana

Job description

Key Responsibilities
Model & AI Agent Deployment

Own end-to-end deployment of ML models and AI agents into production environments.

Manage model versioning, rollback strategies, and lifecycle management.

Ensure high availability, scalability, and reliability of deployed systems.

CI/CD for AI Systems

Design and implement CI/CD pipelines tailored for ML and GenAI workflows.

Automate model training, testing, validation, and deployment processes.

Integrate model testing frameworks, performance checks, and compliance gates into pipelines.

Enable seamless integration between development, staging, and production environments.

Infrastructure & GPU Optimization

Manage and optimize GPU-based workloads for model training and inference.

Monitor and improve compute utilization, cost efficiency, and latency.

Administer and maintain cloud (AWS/Azure/GCP) and/or on-prem infrastructure environments.

Support containerized deployments using Docker and Kubernetes.

Performance Monitoring & Observability

Implement infrastructure and model performance monitoring systems.

Track system health, latency, throughput, resource utilization, and failure rates.

Establish alerting, logging, and incident response processes.

Continuously improve system performance and reliability through proactive monitoring.

Vector Database & Data Infrastructure Management

Deploy and manage vector databases (e.g., FAISS, Pinecone, Weaviate, Chroma).

Optimize indexing, embedding pipelines, and retrieval performance for GenAI applications.

Ensure high availability and backup strategies for AI data systems.

Security, Access Control & Compliance

Implement role-based access control (RBAC) and secure authentication mechanisms.

Ensure infrastructure and AI systems comply with pharma regulatory standards and internal governance policies.

Manage sensitive data securely, including encryption (at rest and in transit).

Support audit readiness and documentation for compliance reviews.

Technical Skills & Competencies
  • Strong programming skills in Python and scripting languages
  • Hands-on experience with CI/CD tools (Jenkins, GitHub Actions, GitLab CI, Azure DevOps,etc.)
  • Experience with Docker, Kubernetes, Helm
  • Knowledge of ML lifecycle tools (MLflow, Kubeflow, Airflow, etc.)
  • Experience with GPU optimization and distributed training frameworks
  • Familiarity with cloud platforms (AWS/Azure/GCP) and hybrid infrastructure
  • Experience managing vector databases
  • Knowledge of monitoring tools (Prometheus, Grafana, ELK, Datadog, etc.)
  • Understanding of data security, encryption, compliance frameworks (GxP, HIPAA preferred)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Operations Engineer
AI Operations Engineer

Shashwath Solution • Pune District

On-site
INR 1,500,000 - 2,100,000
AI Operations Engineer
AI Operations Engineer

Shashwath Solution • Dadri

Hybrid
INR 1,500,000 - 2,300,000
Senior AI/ML Engineer
Senior AI/ML Engineer

IntraEdge • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Senior AI Engineer
Senior AI Engineer

Genzeon Corporation • Pune District

On-site
INR 2,000,000 - 5,000,000
Senior AI Engineer
Senior AI Engineer

Genzeon Global • Pune District

On-site
INR 2,500,000 - 5,000,000
Senior AI/ML Engineer
Senior AI/ML Engineer

NB Healthcare Technologies • Hyderabad

Hybrid
INR 1,800,000 - 3,000,000
AI Team Lead
AI Team Lead

Genzeon Global • Hyderabad

On-site
INR 4,500,000 - 7,500,000
Machine Learning / GenAI Engineer
Machine Learning / GenAI Engineer

HCLTech • Dadri

On-site
INR 1,500,000 - 2,600,000
Senior Data Scientist
Senior Data Scientist

Neurealm • Gurugram District

On-site
INR 900,000 - 1,300,000
Machine Learning Engineer
Machine Learning Engineer

Amgen SA • Hyderabad

On-site
INR 1,400,000 - 2,200,000