Machine Learning Operations Site Reliability Engineer

Persistent

Bengaluru

Hybrid

INR 2,800,000 - 4,200,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive salary and benefits
Culture focused on talent development
Cutting-edge technologies
Employee engagement initiatives
Annual health check-ups
Insurance coverage: group term life,个人

Job summary

Persistent is seeking a skilled Machine Learning Operations Site Reliability Engineer to join its AI and Engineering team in Bengaluru. You will design, implement, and manage MLOps and LLMOps platforms across the ML lifecycle from development to production deployment, monitoring, and optimization.

The role emphasizes CI/CD, training pipelines, observability, security, and collaboration with data scientists, ML engineers, and platform architects to deliver reliable AI solutions in a fast-paced

Qualifications

  • 5-8 years of experience in Machine Learning Operations, Site Reliability Engineering, DevOps, Platform Engineering, or related domains.
  • Strong hands-on experience with MLOps and Machine Learning lifecycle management.
  • Experience designing and implementing scalable AI/ML deployment pipelines.
  • Strong understanding of LLMOps concepts and deployment of Generative AI applications.
  • Experience supporting production AI/ML platforms and mission-critical workloads.
  • Knowledge of model versioning, experiment tracking, model registry, and governance frameworks.
  • Experience implementing observability, logging, monitoring, tracing, and alerting solutions for AI applications.
  • Strong understanding of security, compliance, and responsible AI practices.
  • Experience with performance benchmarking, model evaluation, and optimization techniques.
  • Hands-on expertise with automation frameworks and infrastructure management.
  • Strong programming experience in Python and automation scripting.
  • Experience with containerization technologies such as Docker.
  • Hands-on experience with Kubernetes and container orchestration platforms.
  • Familiarity with cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Understanding of CI/CD practices, Infrastructure as Code, and DevOps methodologies.
  • Experience deploying and managing RAG-based and Agentic AI applications.
  • Knowledge of multimodal AI workflows and Large Language Models.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Excellent communication and collaboration skills with the ability to work across cross-functional teams.
  • Ability to operate effectively in fast-paced and dynamic enterprise environments.

Responsibilities

  • Design, implement, and maintain enterprise-grade MLOps and LLMOps pipelines for AI and machine learning workloads.
  • Build secure, scalable, and reproducible model development, testing, deployment, and monitoring frameworks.
  • Support the operationalization of machine learning models and Large Language Model (LLM)-based applications.
  • Develop and manage Agentic AI and Retrieval-Augmented Generation (RAG) applications in production environments.
  • Implement CI/CD and CT (Continuous Training) pipelines for machine learning models and AI solutions.
  • Establish observability frameworks for monitoring model accuracy, latency, drift, performance, availability, and reliability.
  • Conduct model benchmarking, load testing, and performance optimization activities.
  • Implement security controls, governance standards, and compliance requirements across AI platforms.
  • Automate infrastructure provisioning, deployment, monitoring, and operational workflows.
  • Support application lifecycle management across development, testing, staging, and production environments.
  • Monitor platform health and proactively identify potential risks, bottlenecks, and reliability issues.
  • Troubleshoot production incidents and participate in root cause analysis activities.
  • Create and maintain technical documentation, architecture guidelines, operational runbooks, code samples, and implementation blueprints.
  • Collaborate with Data Scientists, ML Engineers, Security Teams, and Platform Architects to deliver reliable AI solutions.
  • Drive platform improvements focused on scalability, resiliency, operational efficiency, and automation.
  • Implement best practices for model lifecycle management, versioning, governance, and auditability.

Skills

MLOps
Site Reliability Engineering
DevOps
Platform Engineering
Python
Automation scripting
Docker
Kubernetes
Cloud platforms (AWS/Azure/GCP)
CI/CD
Infrastructure as Code
LLMOps
Agentic AI
RAG applications

Tools

Docker
Kubernetes
Terraform
Git

Job description

About Position:

We are seeking a skilled Machine Learning Operations Site Reliability Engineer (SRE) to join our growing AI and Engineering team. In this role, you will be responsible for designing, implementing, and managing robust MLOps and LLMOps platforms that support the complete machine learning lifecycle, from model development to production deployment, monitoring, and optimization.

  • Role: Machine Learning Operations Site Reliability Engineer
  • Location: Bangalore
  • Experience: 5 to 8 Years
  • Job Type: Full-Time Employment
What You'll Do:
  • Design, implement, and maintain enterprise-grade MLOps and LLMOps pipelines for AI and machine learning workloads.
  • Build secure, scalable, and reproducible model development, testing, deployment, and monitoring frameworks.
  • Support the operationalization of machine learning models and Large Language Model (LLM)-based applications.
  • Develop and manage Agentic AI and Retrieval-Augmented Generation (RAG) applications in production environments.
  • Implement CI/CD and CT (Continuous Training) pipelines for machine learning models and AI solutions.
  • Establish observability frameworks for monitoring model accuracy, latency, drift, performance, availability, and reliability.
  • Conduct model benchmarking, load testing, and performance optimization activities.
  • Implement security controls, governance standards, and compliance requirements across AI platforms.
  • Automate infrastructure provisioning, deployment, monitoring, and operational workflows.
  • Support application lifecycle management across development, testing, staging, and production environments.
  • Monitor platform health and proactively identify potential risks, bottlenecks, and reliability issues.
  • Troubleshoot production incidents and participate in root cause analysis activities.
  • Create and maintain technical documentation, architecture guidelines, operational runbooks, code samples, and implementation blueprints.
  • Collaborate with Data Scientists, ML Engineers, Security Teams, and Platform Architects to deliver reliable AI solutions.
  • Drive platform improvements focused on scalability, resiliency, operational efficiency, and automation.
  • Implement best practices for model lifecycle management, versioning, governance, and auditability.
Expertise You'll Bring:
  • 5-8 years of experience in Machine Learning Operations, Site Reliability Engineering, DevOps, Platform Engineering, or related domains.
  • Strong hands-on experience with MLOps and Machine Learning lifecycle management.
  • Experience designing and implementing scalable AI/ML deployment pipelines.
  • Strong understanding of LLMOps concepts and deployment of Generative AI applications.
  • Experience supporting production AI/ML platforms and mission-critical workloads.
  • Knowledge of model versioning, experiment tracking, model registry, and governance frameworks.
  • Experience implementing observability, logging, monitoring, tracing, and alerting solutions for AI applications.
  • Strong understanding of security, compliance, and responsible AI practices.
  • Experience with performance benchmarking, model evaluation, and optimization techniques.
  • Hands-on expertise with automation frameworks and infrastructure management.
  • Strong programming experience in Python and automation scripting.
  • Experience with containerization technologies such as Docker.
  • Hands-on experience with Kubernetes and container orchestration platforms.
  • Familiarity with cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Understanding of CI/CD practices, Infrastructure as Code, and DevOps methodologies.
  • Experience deploying and managing RAG-based and Agentic AI applications.
  • Knowledge of multimodal AI workflows and Large Language Models.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Excellent communication and collaboration skills with the ability to work across cross-functional teams.
  • Ability to operate effectively in fast-paced and dynamic enterprise environments.
Benefits:
  • Competitive salary and benefits package
  • Culture focused on talent development with quarterly growth opportunities and company-sponsored higher education and certifications
  • Opportunity to work with cutting-edge technologies
  • Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards
  • Annual health check-ups
  • Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents
Values-Driven, People-Centric & Inclusive Work Environment:

Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds.

  • We support hybrid work and flexible hours to fit diverse lifestyles.
  • Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities.
  • If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment

"Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind."

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Operations Site Reliability Engineer
Machine Learning Operations Site Reliability Engineer

Persistent Systems Limited • Bengaluru

Hybrid
INR 1,800,000 - 2,400,000
Hybrid work arrangement
Education sponsorship & growth
Healthcare & life insurance
AI/ML Engineer
AI/ML Engineer

Persistent • Hyderabad, Pune District

Hybrid
INR 3,500,000 - 5,200,000
Competitive salary
Talent development
Flexible hours
+2
AI/ML Engineer
AI/ML Engineer

Persistent Systems • Pune District

On-site
INR 4,000,000 - 6,000,000
Flexible work hours
Long Service awards
Health insurance
Sr. ML Engineer
Sr. ML Engineer

Persistent • Pune District

Hybrid
INR 3,500,000 - 7,000,000
Hybrid work
Competitive salary
Education sponsorship
+3
Senior Java Developer(AI/ML)
Senior Java Developer(AI/ML)

Persistent • Pune District

On-site
INR 2,500,000 - 4,500,000
Competitive salary
Talent development opportunities
Cutting-edge technologies
+4
Senior Java Developer(AI/ML)
Senior Java Developer(AI/ML)

Persistent • Pune District

On-site
INR 2,500,000 - 4,500,000
Competitive salary
Talent development opportunities
Cutting-edge technologies
+4
Machine Learning Specialist
Machine Learning Specialist

Persistent Systems • Pune District

Hybrid
INR 1,500,000 - 2,500,000
Competitive salary
Quarterly growth opportunities
Company-sponsored higher education
+3
Machine Learning Engineer
Machine Learning Engineer

Persistent • Pune District

On-site
INR 2,600,000 - 4,200,000
Competitive salary
Education sponsorships
Flexible work hours
+2
AI/ML MLOps Engineer
AI/ML MLOps Engineer

Persistent Systems • Pune District

Hybrid
INR 3,000,000 - 5,400,000
Competitive salary and benefits
Growth opportunities and sponsoredcert
Employee engagement initiatives
+1
Sr. ML Engineer
Sr. ML Engineer

Persistent Systems • Pune District

Hybrid
INR 1,800,000 - 3,500,000
Hybrid work
Long Service awards
Insurance coverage