AI/MLOps SRE Lead Engineer

Regeneron

Hyderabad

Hybrid

INR 2,500,000 - 4,500,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Regeneron is seeking an AI-MLOps SRE Lead Engineer to drive reliability and scalability across our AI, ML, and cloud platforms. The Hyderabad-based hybrid role demands leadership in SRE, MLOps, and platform modernization, collaborating with Data Science, AI Engineering, and Platform teams.

You will design infrastructure as code, CI/CD pipelines, and enterprise-grade observability, enabling fast and safe AI innovation at scale.

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, AI, or related field.
  • 6–8 years of experience in Site Reliability/Platform Engineering or DevOps with enterprise-scale delivery.
  • Hands-on across two+ public clouds (AWS, GCP, or Azure).
  • Deep ML platform knowledge including Databricks, SageMaker, Dataiku, Vertex AI.
  • End-to-end ML workflows: training, deployment, monitoring, orchestration.
  • Terraform, Pulumi, AWS CDK, and modern CI/CD practices.
  • Proficient in Python, Go, or Bash.
  • Experience with observability: Prometheus, Datadog, OpenTelemetry, distributed tracing.

Responsibilities

  • Drive reliability, scalability, observability in multi-cloud environments.
  • Design and operate resilient ML platform infrastructure for LLMs and AI Agents.
  • Lead CI/CD, IaC, self-healing systems, and platform automation.
  • Develop enterprise ChatOps and AI-enabled operational workflows.
  • Mentor engineers and influence reliability and MLOps strategy.

Skills

SRE mindset
Distributed systems
Cloud architecture
Python/Go scripting

Education

Bachelor's in CS/IT/Engineering

Tools

Databricks
SageMaker
Dataiku
Vertex AI
Terraform
Kubernetes (EKS/GKE/AKS)
Prometheus/Grafana
OpenTelemetry

Job description

Build our future together

Regeneron is founded on the belief that the right idea, combined with the right team, can lead to significant transformations. Our growing global network is dedicated to inventing, developing, and commercializing medicines that change lives for those with serious diseases. In doing so, we are pioneering innovative approaches to science, manufacturing, and commercialization, as well as redefining our understanding of health.

At Regeneron Digital & Technology, we are expanding our AI and Platform Engineering capabilities to support next-generation intelligent systems, machine learning platforms, and cloud-native technologies. We are seeking an AI-MLOps SRE Lead Engineer to drive reliability, scalability, observability, and operational excellence across our AI, ML, and cloud ecosystem. This role will lead the design and operation of resilient platforms supporting machine learning workloads, LLMs, AI Agents, and enterprise-scale automation while enabling engineering teams to innovate with speed and confidence.

When & Where

Hyderabad (Hybrid)

Discover your role
  • Drive service reliability, availability, and performance across multi-cloud environments, establishing SLOs, SLIs, error budgets, and reliability standard methodologies.
  • Design, build, and scale enterprise ML platform infrastructure using technologies such as Dataiku, Amazon SageMaker AI, Databricks, and Google Vertex AI.
  • Develop AI-driven observability capabilities using anomaly detection, predictive analytics, and automated remediation solutions to proactively identify and resolve operational issues.
  • Lead the implementation and monitoring of LLM, SLM, RAG, and AI Agent platforms, ensuring performance, governance, operational efficiency, and scalability.
  • Design and implement Infrastructure as Code, CI/CD pipelines, self-healing systems, and platform automation capabilities to improve engineering productivity and operational resilience.
  • Architect enterprise ChatOps solutions that integrate operational events, AI workflows, observability platforms, and automated remediation capabilities.
  • Partner with Data Science, AI Engineering, and Platform teams to deliver secure, scalable, and production-ready AI/ML solutions.
  • Evaluate emerging AI-native operational technologies and integrate innovative solutions that enhance platform reliability, engineering efficiency, and business value.
  • Conduct technical debt assessments, identify architectural risks, and provide strategic recommendations to improve enterprise platform maturity.
  • Serve as a technical leader and trusted advisor, mentoring engineers and influencing reliability engineering, MLOps, and cloud platform strategy across the organization.
This role requires
  • Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, Artificial Intelligence, or a related subject area; Master's degree preferred.
  • 6-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related technology fields with enterprise-scale delivery experience.
  • Strong hands-on experience operating across two or more major cloud platforms, including AWS, GCP, or Azure.
  • Deep expertise with ML platform technologies including Databricks, Amazon SageMaker AI, Dataiku, and Google Vertex AI.
  • Proven experience implementing end-to-end ML workflows including model training, deployment, experiment tracking, monitoring, and pipeline orchestration.
  • Advanced proficiency in Infrastructure as Code tools such as Terraform, Pulumi, AWS CDK, and modern CI/CD automation practices.
  • Strong programming and scripting skills in Python, Go, Bash, or similar languages.
  • Experience building enterprise observability solutions using Prometheus, Grafana, Datadog, OpenTelemetry, distributed tracing, metrics, and logging platforms.
  • Demonstrated expertise in anomaly detection, predictive analytics, automated remediation, and AI-assisted operational capabilities.
  • Proven experience designing and implementing enterprise ChatOps solutions and AI-enabled operational workflows.
  • Strong ability to assess technical debt, influence technical strategy, and drive platform modernization initiatives.
  • Experience using AI tools, LLM-powered assistants, and AI Agents to enhance engineering operations and productivity.
  • Experience with Kubernetes and container orchestration platforms such as EKS, GKE, or AKS preferred.
  • Familiarity with MLOps technologies including Kubeflow, Feast, MLflow, and model evaluation frameworks such as LangSmith, RAGAS, Evidently AI, or Weights & Biases preferred.
  • Knowledge of cloud cost optimization, policy-as-code, compliance automation, and multi-cloud governance practices preferred.

Regeneron is an equal opportunity employer and all qualified applicants will receive consideration for employment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/MLOps SRE Lead Engineer
AI/MLOps SRE Lead Engineer

Regeneron Pharmaceuticals, Inc • Hyderabad

Hybrid
INR 4,000,000 - 7,000,000
AI/MLOps SRE Lead Engineer
AI/MLOps SRE Lead Engineer

Regeneron India Private Limited • Hyderabad

Hybrid
INR 4,000,000 - 7,500,000
Hybrid work model
Comprehensive benefits
Senior Director, Technology Operations & Platform Reliability
Senior Director, Technology Operations & Platform Reliability

Initial Therapeutics, Inc. • India

Hybrid
INR 4,000,000 - 7,000,000
Senior Director - AI, Data & Solutions
Senior Director - AI, Data & Solutions

Regeneron • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior Director AI, Data & Solutions
Senior Director AI, Data & Solutions

Regeneron • Hyderabad

Hybrid
INR 5,000,000 - 9,000,000
Performance bonuses
Health & wellness benefits
Pension or retirement benefits
Senior Director AI, Data & Solutions
Senior Director AI, Data & Solutions

Initial Therapeutics, Inc. • India

Hybrid
INR 5,000,000 - 8,000,000
Senior Director, Technology Operations & Platform Reliability
Senior Director, Technology Operations & Platform Reliability

Regeneron • Hyderabad

Hybrid
INR 4,000,000 - 7,000,000
Staff Ml Ops Engineer [T500-28545]
Staff Ml Ops Engineer [T500-28545]

ANSR • Bengaluru

On-site
INR 4,500,000 - 7,000,000
Associate Machine Learning Engineer
Associate Machine Learning Engineer

Amgen Inc. (IR) • Hyderabad

On-site
INR 800,000 - 1,200,000
Senior Director AI, Data & Solutions
Senior Director AI, Data & Solutions

BioSpace • Hyderabad

Hybrid
INR 3,000,000 - 6,000,000