MLOps Engineer

Solve It Consultant

Gurugram District, Chennai District

On-site

INR 2,500,000 - 4,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Solve It Consultant is hiring a Lead/Principal MLOps Engineer to own end-to-end ML infrastructure from Gurugram or Chennai, onsite five days a week. You will drive scalable, reliable ML pipelines, feature stores, and model deployments, collaborating with Data Science, Data Engineering, and DevOps teams.

The role demands 8+ years of engineering experience, strong cloud and container orchestration skills, and a proactive approach to performance, security, and cost management.

Qualifications

  • 8+ years of professional engineering experience with 4+ years in MLOps at scale.
  • Bachelor’s or Master’s degree in CS/SE/IT or related field.
  • Hands-on with ML platforms, CI/CD, IaC and multi-cloud environments.

Responsibilities

  • Own production ML infrastructure across multi-cloud or hybrid environments.
  • Automate CI/CD pipelines and zero-downtime model releases.
  • Monitor, diagnose, and resolve pipeline and data issues in production.
  • Ensure high availability, DR readiness, and cost optimization.

Skills

MLOps
CI/CD
Docker
Kubernetes
Terraform
Python
Monitoring
Cloud (AWS/Azure/GCP)
Security governance

Education

Bachelor’s or Master’s in CS/Software/IT

Tools

MLflow
Kubeflow
Airflow
Weights & Biases
SageMaker
Prometheus
Grafana
ELK/Datadog
Arize/Whylogs/Evidently AI

Job description

Job Title: Lead / Principal MLOps Engineer

Job Location: Gurugram or Chennai (5 Days Work From Office)

Experience Level: 8+ Years

Employment Type: Full-time

Department: Artificial Intelligence / Cloud Engineering & Infrastructure

About the Role

We are seeking an experienced and battle-tested MLOps Engineer with 8+ years of expertise to own, scale, and maintain our end-to-end Machine Learning and Generative AI infrastructure. In this role, you will be responsible for ensuring the high availability, reliability, and continuous performance of production ML pipelines, automated monitoring systems, and model deployment workflows.

Working on-site from our Gurugram or Chennai offices, you will bridge the gap between Data Science, Data Engineering, and DevOps building resilient infrastructure that powers production AI systems at scale.

Key Responsibilities
1. Production ML Infrastructure & Platform Ownership
  • Architect, deploy, and manage robust, scalable Machine Learning infrastructure across multi-cloud or hybrid environments (AWS, Azure, or GCP).
  • Automate CI/CD pipelines for ML models (CT/CD), ensuring seamless model deployment, rollback strategies, and zero-downtime releases.
  • Design and maintain feature stores, model registries, and containerized deployment runtime environments (Docker, Kubernetes/KServe).
2. Pipeline Monitoring & Operational Issue Resolution
  • Establish real-time telemetry, monitoring, and alerting frameworks to track system health, inference latency, GPU/CPU utilization, and pipeline throughput.
  • Serve as the primary escalation point for production operational incidents rapidly diagnosing and resolving pipeline failures, data drift, and infrastructure bottlenecks.
  • Conduct root-cause analysis (RCA) for operational outages and implement permanent remediations to safeguard system SLAs.
3. System Reliability & Performance Optimization
  • Ensure strict adherence to high-availability (99.9%+ uptime), fault tolerance, and disaster recovery standards across all ML workflows.
  • Monitor models in production for concept drift, data drift, and latency degradation, triggering automated retraining and re-deployment workflows.
  • Optimize resource allocation, cluster auto-scaling, and compute usage to reduce cloud infrastructure costs without compromising speed or reliability.
4. Configuration Management & Security governance
  • Implement and manage Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or Ansible to enforce configuration consistency across environments.
  • Manage configuration updates, version control, and secret management for complex, distributed ML and LLM microservices.
  • Partner with InfoSec teams to enforce data governance, access controls, compliance standards, and security patches across the ML ecosystem.
Requirements & Qualifications
  • Experience: 8+ years of professional engineering experience, with at least 4+ years dedicated to MLOps, ML Platform Infrastructure, or Cloud Engineering at scale.
  • Education: Bachelor’s or Master’s degree in Computer Science, Software Engineering, Information Technology, or a related field.
  • Core Technical Expertise:

MLOps Frameworks: Hands-on mastery with platforms like MLflow, Kubeflow, Airflow, Weights & Biases, Argo Workflows, or SageMaker.

Containerization & Orchestration: Advanced expertise in Docker, Kubernetes (EKS/GKE/AKS), Helm charts, and ingress controllers.

CI/CD & IaC: Strong command of GitHub Actions, GitLab CI, Jenkins, and Infrastructure-as-Code tools like Terraform.

Monitoring & Observability: Proficiency with Prometheus, Grafana, ELK Stack, Datadog, or specialized ML monitoring tools (Evidently AI, Whylogs, Arize).

Programming & Scripting: Expert proficiency in Python, Bash, and SQL for automation, CLI tooling, and service integration.

  • Location & Work Mode: Willingness to work 5 days from office at either our Gurugram or Chennai location.
Preferred Qualifications
  • Experience with LLMOps (deploying, serving, and monitoring Large Language Models using vLLM, Ollama, or Triton Inference Server).
  • Hands-on experience managing GPU compute clusters, CUDA acceleration, and distributed inference/training.
  • Relevant certifications in AWS/GCP/Azure Cloud Architecture or Kubernetes (CKA/CKAD).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

GyanSys Inc. • Bengaluru

Hybrid
INR 1,400,000 - 2,100,000
MLOps Engineer
MLOps Engineer

EAZYGURUS IT TRAINING • Hyderabad

On-site
INR 1,100,000 - 1,700,000
MLOps Engineer
MLOps Engineer

eazygurus • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Sr. Manager, Enterprise Systems
Sr. Manager, Enterprise Systems

Skyworks Solutions, Inc. • Bengaluru

On-site
Competitive salary
Career growth opportunities
Referral bonus program of Rs200,000
MLops Engineer
MLops Engineer

Straive • Bengaluru Urban

On-site
INR 1,000,000 - 1,700,000
MLOps Engineer
MLOps Engineer

In Time Tec • Jaipur, Bengaluru

On-site
INR 1,500,000 - 2,500,000
MLOps Engineer
MLOps Engineer

Aspyra Hr Services • Dadri, Gurugram District, Delhi

Hybrid
INR 1,200,000 - 2,400,000
ML Engineer
ML Engineer

Aligned Automation Services • Pune District

On-site
INR 167,400 - 279,000
MLOps Manager
MLOps Manager

Anblicks • Hyderabad

On-site
INR 2,000,000 - 3,000,000
MLOps Platform Engineer (Chennai / Pune)
MLOps Platform Engineer (Chennai / Pune)

Money Forward India • Chennai District

On-site
INR 3,500,000 - 6,000,000