SRE Engineer

Siemens Digital Industries Software

Maharashtra

Hybrid

INR 2,400,000 - 3,600,000

Full time

37 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary
Hybrid work model
Career growth opportunities

Job summary

Siemens Digital Industries Software is seeking an experienced DevOps/MLOps engineer to define, monitor, and enforce SLOs/SLIs for AI services and to ensure scalable, fault-tolerant infrastructure for model serving and data workflows.

You will manage Azure cloud resources with IaC tools (Terraform/ARM), oversee Kubernetes clusters (Docker, AKS), and build observability with Prometheus, Grafana, Datadog, and Open Telemetry, while leading incident response and RCA processes.

Qualifications

  • Enforce SLOs/SLIs and error budgets for AI services.
  • Lead incident response, RCA, and post-mortem processes to minimize downtime and recurrence.
  • Proactively identify and resolve reliability risks before they impact end users.
  • Maintain scalable, fault-tolerant infrastructure for data pipelines, ML workflows and data pipelines.
  • Handle Azure cloud infrastructure using Terraform/ARM templates.
  • Oversee Kubernetes clusters and containerized workloads (Docker) for AI microservices.
  • Maintain observability stacks with Prometheus, Grafana, Datadog, or OpenTelemetry.
  • Develop alerting systems to detect anomalies in AI model performance and latency.
  • Automate operational tasks with scripting (Python, Bash, Go).
  • Build and maintain CI/CD pipelines for model deployment and service updates.
  • Implement MLOps practices to streamline deployment and monitoring of ML models.
  • Ensure infrastructure aligns with security guidelines and regulatory standards.
  • Collaborate with data engineers and data scientists to productionize AI systems.
  • Foster a reliability-first culture across engineering teams.
  • Contribute to on-call runbooks and playbooks.

Responsibilities

  • Enforce SLOs, SLIs, and error budgets for AI services
  • Lead incident response, RCA, and post-mortem processes to minimize downtime and prevent recurrence
  • Proactively identify and resolve reliability risks before they impact end users
  • Maintain scalable, fault-tolerant infrastructure for data pipelines, ML workflows and data workflows
  • Handle cloud infrastructure (Azure) using Infrastructure as Code (IaC) tools such as Terraform
  • Oversee Kubernetes clusters and containerized workloads (Docker) supporting AI microservices
  • Maintain comprehensive observability stacks (metrics, logs, traces) using tools like Prometheus, Grafana, Datadog, or Open Telemetry
  • Develop intelligent alerting systems to detect anomalies in AI model performance, latency, and efficiency
  • Automate repetitive operational tasks through scripting and tooling (Python, Bash, Go)
  • Build and maintain CI/CD pipelines for model deployment and service updates
  • Implement MLOps practices to streamline the deployment and monitoring of ML models in production
  • Ensure infrastructure and services align with security guidelines and relevant regulatory standards
  • Partner with data engineers and data scientists to make AI systems production-ready
  • Foster a reliability-first culture across engineering teams
  • Contribute to on-call rotations and continuously improve on-call runbooks and playbooks

Skills

Problem-solving
Analytical thinking
Proactive issue resolution
Security awareness
Multi-tenant architecture
Cloud design

Education

Bachelor's degree in CS/IT

Tools

Azure AKS
Azure DevOps
Azure ARM
Terraform
Ansible
Jenkins
GitHub Actions
Prometheus
Grafana
Datadog
OpenTelemetry
Docker
Kubernetes
ELK Stack

Job description

Siemens Digital Industries Software is a leading provider of solutions for the design, simulation, and manufacture of products across many different industries. Formula 1 cars, skyscrapers, ships, space exploration vehicles, and many of the objects we see in our daily lives are being conceived and manufactured using our Product Lifecycle Management (PLM) software.

Define, monitor, and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets for AI services, understand scalable, fault-tolerant infrastructure for AI model serving, training pipelines, and data workflows, maintain comprehensive observability stacks (metrics, logs, traces) using tools like Datadog or Open Telemetry.

Responsibilities
  • Enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets for AI services
  • Lead incident response, root cause analysis (RCA), and post-mortem processes to minimize downtime and prevent recurrence
  • Proactively identify and resolve reliability risks before they impact end users
  • Maintain scalable, fault-tolerant infrastructure for data pipelines, ML workflows and data workflows
  • Handle cloud infrastructure (Azure) using Infrastructure as Code (IaC) tools such as Terraform
  • Oversee Kubernetes clusters and containerized workloads (Docker) supporting AI microservices
  • Maintain comprehensive observability stacks (metrics, logs, traces) using tools like Prometheus, Grafana, Datadog, or Open Telemetry
  • Develop intelligent alerting systems to detect anomalies in AI model performance, latency, and efficiency
  • Automate repetitive operational tasks through scripting and tooling (Python, Bash, Go)
  • Build and maintain CI/CD pipelines for model deployment and service updates
  • Implement MLOps practices to streamline the deployment and monitoring of ML models in production
  • Ensure infrastructure and services align with security guidelines and relevant regulatory standards
  • Partner with data engineers and data scientists to make AI systems production-ready
  • Foster a reliability-first culture across engineering teams
  • Contribute to on-call rotations and continuously improve on-call runbooks and playbooks
Qualifications
  • Bachelor’s degree in computer science, Information Technology, or a related field with 6-8 years of meaningful experience.
  • Demonstrated experience supporting production-grade, high-availability systems
  • Strong experience with Azure cloud services, including Azure Kubernetes Service (AKS), Azure DevOps, Azure Resource Manager (ARM), and monitoring tools.
  • Proven expertise in handling multi-tenant, microservices-based architectures and deploying high-availability solutions.
  • Proficiency in Infrastructure as Code (IaC) tools like Terraform, Ansible, or ARM templates for handling and automating cloud resources.
  • Hands-on experience with CI/CD tools (e.g., Jenkins, GitHub Actions, Azure DevOps/MLOps) and configuration management.
  • Solid understanding of containerization technologies (e.g., Docker) and orchestration (e.g., Kubernetes).
  • Secondary experience or familiarity with security practices, such as identity and access management (IAM), threat detection, and compliance standards.
  • Solid understanding of secure supply chain practices, including management of code dependencies and open-source libraries.
  • Strong problem-solving and analytical skills with a proactive approach to operational and security issues.
Preferred Skills
  • Experience with supervising and logging tools (e.g., Prometheus, Grafana, ELK Stack) for proactive issue resolution.
  • Familiarity with zero-trust architecture principles
  • Knowledge of industry compliance standards, such as SOC 2, ISO 27001, and GDPR, is a plus.
  • Experience with scripting and automation tools, such as Python, Bash, or PowerShell.
What We Offer
  • Competitive salary and comprehensive benefits package.
  • Opportunities to work on ground breaking projects in cloud infrastructure and MLOps.
  • A collaborative, innovative work environment with room for career growth.
Why us?

Working at Siemens Software means flexibility - Choosing between working at home and the office at other times is the norm here. We offer great benefits and rewards, as you'd expect from a world leader in industrial software.

A collection of over 377,000 minds building the future, one day at a time in over 200 countries. We're dedicated to equality, and we welcome applications that reflect the diversity of the communities we Work in. All employment decisions at Siemens are based on qualifications, merit, and business need. Bring your curiosity and creativity and help us shape tomorrow!

Siemens Software. Transform the Everyday

#SWSaaS

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Engineer
SRE Engineer

Siemens Mobility • Pune District

Hybrid
INR 3,500,000 - 6,500,000
Service Delivery Manager, Enterprise AI
Service Delivery Manager, Enterprise AI

Siemens • Maharashtra

On-site
INR 900,000 - 1,300,000
Senior AI Engineer
Senior AI Engineer

Siemens Energy • Maharashtra

Hybrid
INR 3,000,000 - 6,000,000
Flexible and hybrid working env
Continuous professional development
Collaborative international team
Senior Software Engineer, Simcenter X Platform
Senior Software Engineer, Simcenter X Platform

Siemens Digital Industries Software • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Senior Engineering Manager
Senior Engineering Manager

Siemens Mobility • New Delhi

On-site
INR 4,200,000 - 8,000,000
Software QA Engineer - Advanced, Test Automation & AI
Software QA Engineer - Advanced, Test Automation & AI

Siemens Digital Industries Software • Bengaluru

Hybrid
INR 1,800,000 - 3,500,000
Services Project Manager
Services Project Manager

Siemens Digital Industries Software • Maharashtra

Hybrid
INR 4,523,749 - 6,031,665
Competitive salary and benefits
Global team collaboration
Career development program
+1
Senior Software Engineer Simcenter X Platform
Senior Software Engineer Simcenter X Platform

Siemens Digital Industries Software • Bengaluru

On-site
INR 2,800,000 - 4,000,000
Infrastructure Engineer – Quality Engineering & Test Automation
Infrastructure Engineer – Quality Engineering & Test Automation

Siemens Digital Industries Software • Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
ServiceNow Technical Lead
ServiceNow Technical Lead

Siemens Energy • Gurugram District

Hybrid
INR 1,800,000 - 3,000,000
Remote work up to 2 days/week
Medical benefits
Learn@Siemens-Energy access