Lead Site Reliability Engineer - Observability

SimCorp

Hyderabad

Hybrid

INR 1,500,000 - 2,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Global hybrid work policy
Inclusive and diverse company culture
Work-life balance

Job summary

SimCorp in Hyderabad is seeking a Site Reliability Engineer to enhance mission-critical environments for CloudNative products. Candidates should have over 5 years of experience and a strong command of Azure Cloud and observability frameworks.

This role offers global hybrid work, promoting both a supportive culture and work-life balance. Benefits include an inclusive atmosphere where all contributions are valued.

Qualifications

  • 5+ years experience in Site Reliability, Observability, DevOps, or Cloud Engineering roles.
  • Strong understanding of instrumenting, tracing, and correlating AI/LLM workflows.
  • Experience managing both onboarding projects and live production operations.

Responsibilities

  • Support operational and enhancement of CloudNative products and services.
  • Deploy and manage instrumentation for applications.
  • Collaborate with Agile teams in design discussions.

Skills

Site Reliability
Observability
Azure Cloud
Infrastructure as Code
Monitoring tools

Education

Bachelor’s degree in Computer Science

Tools

AzureMonitor
ApplicationInsights
Terraform
Kubernetes
Docker

Job description

About the Role

SimCorp’s Observability strategy is to deliver a consistent & coherent Observability approach across the full SimCorp One ecosystem, encompassing different technology stacks, products and services across the organization. This approach enables seamless observation of our products and services, efficient investigation of emerging problems, and the delivery of high‑quality software while staying within agreed resolution times. It also provides insights into KPIs, SLOs, SLAs and cost attribution.

Responsibilities
  • Support the operational and enhancement of mission‑critical environments for both new and existing CloudNative products & services.
  • Deploy and manage instrumentation for applications to gain granular insights into service health.
  • Assist engineering teams in implementing and maintaining metrics, logs, and traces for applications & infrastructure.
  • Unify observability tooling across teams, ensuring metrics, logs, and traces flow into a central platform (e.g., ApplicationInsights or equivalent).
  • Enable and configure OpenTelemetry‑based data collection within AzureMonitor ApplicationInsights by leveraging Azure Monitor OpenTelemetry Distro.
  • Make sure AI agent frameworks adopt the semantic convention to ensure interoperability and consistency in observability data.
  • Work with product development teams to enable structured logging, basic distributed tracing, and core metrics.
  • Support incident response by gathering logs, metrics, and traces to perform root‑cause analysis using observability tools.
  • Build tools and automation to eliminate TOIL, improve engineering velocity, developer experience, and system reliability.
  • Define and manage SLOs and error budgets in partnership with Engineering teams.
  • Provide flexible working in regular & evening shift on rotational basis, and provide weekend or on‑call support as needed.
  • Collaborate with Agile teams and take part in design discussions with clients, vendors, and stakeholders.
  • Contribute to knowledge sharing across multiple Product Areas.
  • Leverage a strong foundation in ITIL practices, including problem, change, and incident management.
Qualifications
  • Bachelor’s degree in Computer Science or related field (Master’s is a plus).
  • 5+years experience in Site Reliability, Observability, DevOps, or Cloud Engineering roles.
  • Expertise with Microsoft Azure Cloud.
  • Experience working with observability frameworks like OpenTelemetry and distributed tracing systems.
  • Expertise in Infrastructure as Code (IaC) using Bicep, ARM and Terraform.
  • Strong understanding of instrumenting, tracing, and correlating AI/LLM workflows with infrastructure telemetry.
  • Solid experience in monitoring and logging tools (AzureMonitor, ApplicationInsights, DataDog, Log Analytics).
  • Knowledge of AI/ML‑based anomaly detection, log aggregation and analysis tools such as Microsoft Azure Anomaly Detector or equivalent.
  • Experience with AI/LLM‑based systems (e.g., LangChain, Celery, OpenAI APIs, orchestration frameworks).
  • Experience working with application reliability platforms like Checkly or equivalent.
  • Experience setting up synthetic monitoring using Playwright or equivalent.
  • Solid understanding of networking, containerization (Kubernetes, Docker).
  • Good understanding of APIs, scripting languages like PowerShell, Bash, Kusto and databases (SQL, CosmosDB, PostgresSQL).
  • Familiarity with SimCorp Dimension & Salesforce (User is a plus).
  • Proficiency in IT service management (ITSM) frameworks like ITIL, focusing on incident, change, and problem management to improve operational efficiency.
  • Experience managing both onboarding projects and live production operations.
  • Collaborative mindset and ability to work in cross‑functional teams.
  • Interest in continuous learning and growth within our Product Area.
Benefits
  • Global hybrid work policy – work 2 days a week from the office; remote on remaining days.
  • Inclusive and diverse company culture.
  • Work‑life balance: equilibrium between professional responsibilities and private life.
  • Empowerment: all voices are valued and heard.
  • Career & Growth: individual approach to professional development.

SimCorp is an equal opportunity employer and welcomes applicants from all backgrounds, without regard to race, gender, age, disability or any other protected status under applicable law. We are committed to building a culture where diverse perspectives and expertise are integrated into our everyday work.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

simcorp • Hyderabad

Hybrid
INR 4,000,000 - 6,500,000
Global hybrid work policy
Growth and innovation sprint time
Self-Direction
Lead Software Engineer
Lead Software Engineer

SimCorp • Dadri

Hybrid
INR 1,200,000 - 1,800,000
Competitive salary
Hybrid workplace model
Flexible work hours
+1
Manager, Engineering
Manager, Engineering

SimCorp • Dadri

Hybrid
INR 9,523,000 - 12,381,000
Global hybrid work policy
Growth and innovation opportunities
Inclusivity and diversity focus
+1
Engineering Manager – SaaS Onboarding and Reliability
Engineering Manager – SaaS Onboarding and Reliability

SimCorp • Hyderabad

Hybrid
INR 2,500,000 - 3,500,000
Global hybrid work policy
Personal growth opportunities
Inclusive company culture
+2
Lead/Sr. Backend Software Engineer (C#.NET with Azure)
Lead/Sr. Backend Software Engineer (C#.NET with Azure)

SimCorp • Dadri

On-site
INR 1,200,000 - 1,800,000
Global hybrid work policy - 2 days in office
Inclusive company culture
Focus on work-life balance
+1
Engineering Manager – SaaS Onboarding and Reliability
Engineering Manager – SaaS Onboarding and Reliability

SimCorp • Hyderabad

Hybrid
INR 9,578,000 - 13,410,000
Global hybrid work policy
Growth and innovation opportunities
Inclusive and diverse company culture
Jul_POS23_Software Engineer
Jul_POS23_Software Engineer

SimCorp • Dadri

Hybrid
INR 1,500,000 - 2,800,000
Hybrid workplace model
IP sprints (3 weeks per quarter)
Bonus scheme
+1
Jul_POS24_Software Engineer
Jul_POS24_Software Engineer

SimCorp • Dadri

Hybrid
INR 1,200,000 - 1,800,000
Hybrid workplace model
IP sprints (3 weeks per quarter)
Flexible work hours
+1
Site Reliability Engineer (Application Reliability)
Site Reliability Engineer (Application Reliability)

Alter Domus • Hyderabad

On-site
INR 2,000,000 - 3,200,000
ACCA study leave
Flexible holidays
Mentoring
+3
Jul_POS19_Principal Software Engineer
Jul_POS19_Principal Software Engineer

SimCorp • India

Hybrid
INR 4,000,000 - 7,000,000
Hybrid workplace model
Bonus scheme
Pension
+1