AI-MLOps SRE Lead Engineer

Randstad

Hyderabad

Hybrid

INR 2,400,000 - 3,600,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Randstad in Hyderabad is seeking an AI-MLOps SRE Lead Engineer to drive reliability across multi-cloud environments. You will define SLOs/SLIs, build scalable ML platform infra using Dataiku, SageMaker AI, Databricks and Vertex AI, and develop AI-driven observability with anomaly detection and automated remediation.

You will lead ChatOps, IaC, and CI/CD efforts, partnering with Data Science and Platform teams to deliver secure, production-ready AI/ML solutions.

Qualifications

  • Bachelor's degree in CS/IT/Engineering or related field; Master's preferred.
  • 6–8 years in SRE/Platform/DevOps with enterprise-scale delivery.
  • Hands-on with two or more cloud platforms (AWS, GCP, or Azure).
  • Deep ML platform experience: Databricks, SageMaker, Dataiku, Vertex AI.
  • End-to-end ML workflows: training, deployment, monitoring, orchestration.
  • Proficient in IaC tools Terraform, Pulumi, AWS CDK; modern CI/CD.
  • Strong Python/Go/Bash programming skills.
  • Observability with Prometheus, Grafana, Datadog, OpenTelemetry.
  • Expertise in anomaly detection, predictive analytics, automated remediation.
  • Experience with ChatOps and AI-enabled operational workflows.
  • Ability to assess technical debt and influence platform strategy.
  • Experience using AI tools/agents to boost productivity.
  • Kubernetes experience; EKS/GKE/AKS preferred.
  • Familiar with Kubeflow, Feast, MLflow, LangSmith, RAGAS, Evidently AI, W&B.
  • Knowledge of cost optimization, policy-as-code and multi-cloud governance.

Responsibilities

  • Drive service reliability, availability, and performance across multi-cloud environments.
  • Design, build, and scale enterprise ML platform infrastructure with Dataiku, SageMaker AI, Databricks, Vertex AI.
  • Develop AI-driven observability with anomaly detection, predictive analytics, automated remediation.
  • Lead implementation and monitoring of LLM/SLM/AI Agent platforms for performance and governance.
  • Design IaC, CI/CD pipelines, self-healing systems, and automation to boost productivity.
  • Architect enterprise ChatOps integrating events, AI workflows, observability, remediation.
  • Partner with Data Science, AI Engineering and Platform teams for secure, production-ready AI/ML solutions.
  • Evaluate emerging AI-native ops tech to enhance reliability and business value.
  • Conduct technical debt assessments and provide strategic modernization recommendations.
  • Serve as technical leader and advisor, mentoring engineers in reliability and cloud strategy.

Skills

Python
Go
Bash
Cloud platforms (AWS/GCP/Azure)
Kubernetes
CI/CD automation

Education

Bachelor's degree in Computer Science, IT, Engineering, Data Science or related
Master's degree preferred

Tools

Terraform
Pulumi
AWS CDK
Databricks
SageMaker
Dataiku
Vertex AI
Kubeflow
Prometheus
Grafana
Datadog
OpenTelemetry
EKS
GKE
AKS

Job description

Job Title: AI-MLOps SRE Lead Engineer

Location: Hyderabad. (Hybrid)

Discover your role
  • Drive service reliability, availability, and performance across multi-cloud environments, establishing SLOs, SLIs, error budgets, and reliability standard methodologies.
  • Design, build, and scale enterprise ML platform infrastructure using technologies such as Dataiku, Amazon SageMaker AI, Databricks, and Google Vertex AI.
  • Develop AI-driven observability capabilities using anomaly detection, predictive analytics, and automated remediation solutions to proactively identify and resolve operational issues.
  • Lead the implementation and monitoring of LLM, SLM, RAG, and AI Agent platforms, ensuring performance, governance, operational efficiency, and scalability.
  • Design and implement Infrastructure as Code, CI/CD pipelines, self-healing systems, and platform automation capabilities to improve engineering productivity and operational resilience.
  • Architect enterprise ChatOps solutions that integrate operational events, AI workflows, observability platforms, and automated remediation capabilities.
  • Partner with Data Science, AI Engineering, and Platform teams to deliver secure, scalable, and production-ready AI/ML solutions.
  • Evaluate emerging AI-native operational technologies and integrate innovative solutions that enhance platform reliability, engineering efficiency, and business value.
  • Conduct technical debt assessments, identify architectural risks, and provide strategic recommendations to improve enterprise platform maturity.
  • Serve as a technical leader and trusted advisor, mentoring engineers and influencing reliability engineering, MLOps, and cloud platform strategy across the organization.
This role requires
  • Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, Artificial Intelligence, or a related subject area; Master's degree preferred.
  • 6-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related technology fields with enterprise-scale delivery experience.
  • Strong hands-on experience operating across two or more major cloud platforms, including AWS, GCP, or Azure.
  • Deep expertise with ML platform technologies including Databricks, Amazon SageMaker AI, Dataiku, and Google Vertex AI.
  • Proven experience implementing end-to-end ML workflows including model training, deployment, experiment tracking, monitoring, and pipeline orchestration.
  • Advanced proficiency in Infrastructure as Code tools such as Terraform, Pulumi, AWS CDK, and modern CI/CD automation practices.
  • Strong programming and scripting skills in Python, Go, Bash, or similar languages.
  • Experience building enterprise observability solutions using Prometheus, Grafana, Datadog, OpenTelemetry, distributed tracing, metrics, and logging platforms.
  • Demonstrated expertise in anomaly detection, predictive analytics, automated remediation, and AI-assisted operational capabilities.
  • Proven experience designing and implementing enterprise ChatOps solutions and AI-enabled operational workflows.
  • Strong ability to assess technical debt, influence technical strategy, and drive platform modernization initiatives.
  • Experience using AI tools, LLM-powered assistants, and AI Agents to enhance engineering operations and productivity.
  • Experience with Kubernetes and container orchestration platforms such as EKS, GKE, or AKS preferred.
  • Familiarity with MLOps technologies including Kubeflow, Feast, MLflow, and model evaluation frameworks such as LangSmith, RAGAS, Evidently AI, or Weights & Biases preferred.
  • Knowledge of cloud cost optimization, policy-as-code, compliance automation, and multi-cloud governance practices preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps / ML Platform Engineer
MLOps / ML Platform Engineer

Spearsoftech • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Work from office in Hyderabad
MLOps Engineer - Bangalore Location - Hybrid
MLOps Engineer - Bangalore Location - Hybrid

Genpact • Bengaluru, Delhi, New Delhi

Hybrid
INR 3,000,000 - 6,000,000
AI SRE/ AI Site Reliability Engineer
AI SRE/ AI Site Reliability Engineer

Tata Consultancy Services • Bengaluru

On-site
INR 1,800,000 - 2,800,000
ML Engineer
ML Engineer

Polestar Analytics • Kolkata District

On-site
INR 1,800,000 - 3,200,000
MLops Engineer
MLops Engineer

PwC • Hyderabad, Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
ML OPS - Senior Engineer
ML OPS - Senior Engineer

Iris Software • Dadri

On-site
INR 2,800,000 - 4,800,000
AI/MLOps SRE Lead Engineer
AI/MLOps SRE Lead Engineer

Regeneron Pharmaceuticals • Hyderabad

Hybrid
INR 3,500,000 - 6,500,000
Senior AI/ML Engineer
Senior AI/ML Engineer

NationsBenefits • Hyderabad

On-site
INR 2,800,000 - 4,500,000
AI / MLOps PLATFORM ENGINEER
AI / MLOps PLATFORM ENGINEER

Vrinda International • Bengaluru

On-site
INR 3,500,000 - 5,200,000
Sr. Manager, Enterprise Systems
Sr. Manager, Enterprise Systems

Skyworks Solutions, Inc. • Bengaluru

On-site
Competitive salary
Career growth opportunities
Referral bonus program of Rs200,000