SRE (Site Reliability Engineering)

Cognizant

Bengaluru

Hybrid

INR 2,600,000 - 5,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cognizant in Bengaluru is seeking a seasoned Technical Lead to drive site reliability engineering initiatives. You will guide design, implement SRE practices, and ensure stable day-to-day operations of critical Java applications on Azure in a hybrid work model.

You will lead incident management, monitor performance, optimize Linux hosting, and mentor engineers to reduce incident resolution times while balancing cost and reliability.

Qualifications

  • Experience implementing site reliability engineering practices to improve uptime and resilience.
  • Strong monitoring architecture design and implementation experience.
  • Proven incident management and post-incident review capabilities.

Responsibilities

  • Lead end-to-end SRE design and implementation to improve uptime, scalability, and resilience using Azure and supporting technologies.
  • Develop and maintain robust monitoring strategies for visibility into performance and reliability.
  • Oversee Linux-based system configuration and administration for reliable hosting of Java applications.
  • Provide guidance for Java deployment and runtime tuning on Azure to meet performance targets.
  • Coordinate and manage incident response, triage, and communication with stakeholders.
  • Refine incident management processes including runbooks and escalation paths to reduce MTTD/MTTR.
  • Collaborate with development and operations to embed reliability in SDLC and capacity planning.
  • Optimize Azure resource usage balancing cost, reliability, and compliance.
  • Mentor engineers on monitoring, alerts, and log analysis to improve troubleshooting.

Skills

SRE practices
Monitoring expertise
Incident management
Collaboration across teams
Cloud computing

Tools

Linux
Java
Azure Cloud

Job description

Job Summary

This hybrid role is for a seasoned Technical Lead with strong hands on expertise in site reliability engineering (SRE) monitoring linux java incident management and Azure cloud. The role involves guiding technical decisions improving service reliability and ensuring stable day time operations without travel while collaborating across teams to deliver resilient customer facing systems.



Responsibilities


  • Lead end to end design and implementation of site reliability engineering practices to improve service uptime scalability and resilience across critical applications using Azure cloud and supporting technologies.

  • Drive creation and maintenance of robust monitoring strategies that provide deep visibility into application performance infrastructure health and user experience enabling proactive detection of issues.

  • Oversee configuration optimization and administration of linux based systems to ensure secure reliable and efficient hosting environments for java applications and supporting services.

  • Provide technical guidance for java application deployment and runtime tuning on Azure platforms ensuring that services meet performance latency and throughput requirements under varying workloads.

  • Coordinate and manage incident response activities by triaging issues assessing business impact and guiding teams through structured resolution while maintaining clear communication with stakeholders.

  • Implement and refine incident management processes including runbooks escalation paths and post incident reviews to reduce mean time to detect and mean time to resolve for recurring problems.

  • Collaborate with development and operations teams to embed reliability principles into the software delivery lifecycle including capacity planning dependency analysis and operational readiness checks.

  • Optimize Azure resource usage by guiding configuration of compute storage networking and platform services to balance cost efficiency reliability and compliance with enterprise standards.

  • Mentor engineers on best practices in monitoring alert design and log analysis so that alerts are actionable reduce noise and directly support rapid troubleshooting of production issues.

  • Establish and maintain performance baselines and error budgets for key services and use them to guide prioritization of engineering work focused on reliability and user experience improvements.

  • Coordinate day shift operations for production systems ensuring that routine maintenance changes and validations are executed safely with minimal customer impact in the hybrid work model.

  • Partner with security and compliance teams to ensure that reliability solutions for linux java and Azure environments align with organizational policies and regulatory expectations.

  • Contribute to continuous improvement initiatives by analyzing incident trends system metrics and customer feedback to propose targeted engineering actions that enhance overall service quality.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

PwC India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

Infosys • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Questhiring • Gurugram District

Hybrid
INR 1,200,000 - 1,800,000
Assistant Manager - Azure Site Reliability Engineer
Assistant Manager - Azure Site Reliability Engineer

Promaynov Advisory Services Pvt. Ltd • Bengaluru

On-site
INR 1,400,000 - 2,100,000
Azure Site Reliability Engineer (SRE) - SaaS Operations
Azure Site Reliability Engineer (SRE) - SaaS Operations

Zensar • Pune District, Bengaluru

Hybrid
INR 1,800,000 - 2,400,000
Senior Site Reliability Engineer (Azure) - S
Senior Site Reliability Engineer (Azure) - S

Tata Consultancy Services • Kolkata District, Chennai District, Bengaluru

On-site
INR 2,400,000 - 4,200,000
Resilience and Reliability Engineer
Resilience and Reliability Engineer

EY • Pune District, Gurugram District, Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
Senior Associate Site Reliability Engineer
Senior Associate Site Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Hyderabad

On-site
INR 1,400,000 - 2,000,000
Engineering Manager
Engineering Manager

WaferWire Cloud Technologies • Hyderabad

On-site
INR 4,000,000 - 7,000,000