AI/MLOps SRE Lead Engineer

BioSpace

Hyderabad

Hybrid

INR 4,000,000 - 7,000,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Regeneron in Hyderabad, operating on a hybrid model, is seeking an AI-MLOps SRE Lead Engineer to drive reliability and platform excellence for AI/ML workloads. You will design scalable ML infra using Databricks, SageMaker, Dataiku, and Vertex AI, implement IaC/CI/CD, and mentor teams across Data Science and Platform groups.

You will lead observability initiatives, manage LLM/AI Agent platforms, and partner with cloud-native teams to deliver secure, production-ready AI solutions at scale.

Qualifications

  • Bachelor's degree required; Master's preferred.
  • 6-8 years in SRE, Platform Engineering, DevOps, or related fields with enterprise-scale delivery.
  • Deep cloud expertise across two or more major clouds (AWS, GCP, Azure).
  • Expertise in ML platform tech: Databricks, SageMaker, Dataiku, Vertex AI.
  • Proficient with IaC tools (Terraform, Pulumi, AWS CDK) and CI/CD.
  • Strong programming in Python, Go, Bash.

Responsibilities

  • Drive reliability, availability, and performance across multi-cloud environments.
  • Design, build, and scale enterprise ML platform infrastructure.
  • Develop AI-driven observability with anomaly detection and remediation.
  • Lead deployment and monitoring of LLM/AI Agent platforms.
  • Implement IaC, CI/CD, self-healing and platform automation.
  • Architect ChatOps integrations for AI workflows.
  • Collaborate with Data Science and Platform teams to deliver production-ready AI solutions.
  • Evaluate emerging AI-native ops tech and drive modernization.
  • Mentor engineers and influence reliability strategy.

Skills

SRE
MLOps
Kubernetes
Python
Terraform
Databricks
SageMaker
Dataiku
OpenTelemetry
Prometheus

Education

Bachelor's degree in Computer Science, Engineering, Data Science, AI or related field

Tools

AWS
GCP
Azure
Kubeflow
LangSmith
RAGAS
Weights & Biases
MLflow
Vertex AI
Datadog
OpenTelemetry
Prometheus

Job description

Build our future together Regeneron is founded on the belief that the right idea, combined with the right team, can lead to significant transformations. Our growing global network is dedicated to inventing, developing, and commercializing medicines that change lives for those with serious diseases. In doing so, we are pioneering innovative approaches to science, manufacturing, and commercialization, as well as redefining our understanding of health.

Build our future together Regeneron is founded on the belief that the right idea, combined with the right team, can lead to significant transformations. Our growing global network is dedicated to inventing, developing, and commercializing medicines that change lives for those with serious diseases. In doing so, we are pioneering innovative approaches to science, manufacturing, and commercialization, as well as redefining our understanding of health.

At Regeneron Digital & Technology, we are expanding our AI and Platform Engineering capabilities to support next-generation intelligent systems, machine learning platforms, and cloud-native technologies. We are seeking an AI-MLOps SRE Lead Engineer to drive reliability, scalability, observability, and operational excellence across our AI, ML, and cloud ecosystem. This role will lead the design and operation of resilient platforms supporting machine learning workloads, LLMs, AI Agents, and enterprise-scale automation while enabling engineering teams to innovate with speed and confidence.

When & Where
Hyderabad (Hybrid)
Discover your role
  • Drive service reliability, availability, and performance across multi-cloud environments, establishing SLOs, SLIs, error budgets, and reliability standard methodologies.
  • Design, build, and scale enterprise ML platform infrastructure using technologies such as Dataiku, Amazon SageMaker AI, Databricks, and Google Vertex AI.
  • Develop AI-driven observability capabilities using anomaly detection, predictive analytics, and automated remediation solutions to proactively identify and resolve operational issues.
  • Lead the implementation and monitoring of LLM, SLM, RAG, and AI Agent platforms, ensuring performance, governance, operational efficiency, and scalability.
  • Design and implement Infrastructure as Code, CI/CD pipelines, self-healing systems, and platform automation capabilities to improve engineering productivity and operational resilience.
  • Architect enterprise ChatOps solutions that integrate operational events, AI workflows, observability platforms, and automated remediation capabilities.
  • Partner with Data Science, AI Engineering, and Platform teams to deliver secure, scalable, and production-ready AI/ML solutions.
  • Evaluate emerging AI-native operational technologies and integrate innovative solutions that enhance platform reliability, engineering efficiency, and business value.
  • Conduct technical debt assessments, identify architectural risks, and provide strategic recommendations to improve enterprise platform maturity.
  • Serve as a technical leader and trusted advisor, mentoring engineers and influencing reliability engineering, MLOps, and cloud platform strategy across the organization.
This role requires
  • Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, Artificial Intelligence, or a related subject area; Master's degree preferred.
  • 6-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related technology fields with enterprise-scale delivery experience.
  • Strong hands-on experience operating across two or more major cloud platforms, including AWS, GCP, or Azure.
  • Deep expertise with ML platform technologies including Databricks, Amazon SageMaker AI, Dataiku, and Google Vertex AI.
  • Proven experience implementing end-to-end ML workflows including model training, deployment, experiment tracking, monitoring, and pipeline orchestration.
  • Advanced proficiency in Infrastructure as Code tools such as Terraform, Pulumi, AWS CDK, and modern CI/CD automation practices.
  • Strong programming and scripting skills in Python, Go, Bash, or similar languages.
  • Experience building enterprise observability solutions using Prometheus, Grafana, Datadog, OpenTelemetry, distributed tracing, metrics, and logging platforms.
  • Demonstrated expertise in anomaly detection, predictive analytics, automated remediation, and AI-assisted operational capabilities.
  • Proven experience designing and implementing enterprise ChatOps solutions and AI-enabled operational workflows.
  • Strong ability to assess technical debt, influence technical strategy, and drive platform modernization initiatives.
  • Experience using AI tools, LLM-powered assistants, and AI Agents to enhance engineering operations and productivity.
  • Experience with Kubernetes and container orchestration platforms such as EKS, GKE, or AKS preferred.
  • Familiarity with MLOps technologies including Kubeflow, Feast, MLflow, and model evaluation frameworks such as LangSmith, RAGAS, Evidently AI, or Weights & Biases preferred.
  • Knowledge of cloud cost optimization, policy-as-code, compliance automation, and multi-cloud governance practices preferred.

We are committed to building a workplace with an inclusive culture. Regeneron is an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion or belief (or lack thereof), sex, sexual orientation, gender identity or expression, gender reassignment, marital or civil partnership status, civil status, pregnancy or parental status, age, disability, nationality, citizenship status, ethnic or national origin, membership of the Traveler community, familial status, genetic information, military or veteran status, or any other characteristic protected under applicable law. Where required, we will provide reasonable accommodation to applicants with known disabilities or chronic illnesses during the recruitment process, unless such accommodation would impose undue hardship.

Where necessary, we disclose salary ranges for roles in all countries in which we operate. The final offer will be determined within the relevant range based on the country of employment, specific role level, and your skills and experience. In some countries, collective bargaining agreements (CBAs) may apply and influence certain elements of pay or benefits. Regeneron offers a competitive and comprehensive total rewards package which may include, depending on country and role: annual bonuses or other incentive plans, equity awards, pension or retirement benefits, 401(k) company match, health and wellness programs, fitness centers, insurance benefits (e.g. medical, dental, vision, life and disability), paid time off, and family support benefits. For additional information about Regeneron benefits in the U.S., please visit https://careers.regeneron.com/en/working-at-regeneron/total-rewards/. For other locations, additional information will be provided during the recruitment process.

If you have any questions, please speak with your recruiter.

Please be advised that at Regeneron, we believe we do our best work when we are together. For that reason, many roles are required to be performed onsite. Please speak with your recruiter and hiring manager for more information about onsite expectations for your role and location.

As part of the recruitment process, certain background checks may be conducted in accordance with the laws of the country where the position is based. The purpose of such checks is to verify certain information prior to the commencement of employment such as identity, right to work and educational qualifications.

For jobs in Canada: this posting is for an existing position.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/MLOps SRE Lead Engineer
AI/MLOps SRE Lead Engineer

Regeneron Pharmaceuticals • Hyderabad

Hybrid
INR 3,500,000 - 6,500,000
AI/MLOps SRE Lead Engineer
AI/MLOps SRE Lead Engineer

Regeneron Pharmaceuticals, Inc • Hyderabad

Hybrid
INR 4,000,000 - 7,000,000
AI/MLOps SRE Lead Engineer
AI/MLOps SRE Lead Engineer

Regeneron India Private Limited • Hyderabad

Hybrid
INR 4,000,000 - 7,500,000
Hybrid work model
Comprehensive benefits
Senior Director, Technology Operations & Platform Reliability
Senior Director, Technology Operations & Platform Reliability

Regeneron • Hyderabad

Hybrid
INR 4,000,000 - 7,000,000
Senior Director – Research Solutions
Senior Director – Research Solutions

Regeneron Pharmaceuticals, Inc • Hyderabad

Hybrid
INR 5,000,000 - 6,500,000
Competitive total rewards
Hybrid work flexibility
On-site opportunity in Hyderabad
Senior Director Global Development Solutions
Senior Director Global Development Solutions

Regeneron Pharmaceuticals, Inc • Hyderabad

Hybrid
INR 5,500,000 - 9,000,000
Senior Director G&A Solutions
Senior Director G&A Solutions

Regeneron Pharmaceuticals, Inc • Hyderabad

Hybrid
INR 3,500,000 - 7,000,000
Principal AI Engineer
Principal AI Engineer

BioSpace • Mangaluru

On-site
INR 18,182,000 - 23,923,000
Health and wellness programs
401(k) company match
Paid time off and leaves
Director Commercial Solutions
Director Commercial Solutions

Regeneron • Hyderabad

Hybrid
INR 4,500,000 - 7,500,000
Scientist- Precision Medicine Quantitative Biomarker Sciences
Scientist- Precision Medicine Quantitative Biomarker Sciences

Regeneron Pharmaceuticals, Inc • Hyderabad

Hybrid
INR 1,400,000 - 2,800,000