SRE Lead

3across

Bengaluru

Hybrid

INR 1,500,000 - 2,300,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

3across is seeking an experienced Site Reliability Engineer (SRE) Lead to own critical production support and drive incident management. The role focuses on Python scripting, Linux/Unix, SQL, and automation to improve reliability.

The candidate will lead cross-functional teams during Sev incidents, perform RCA, and develop playbooks and monitoring strategies to prevent recurrence. Bangalore-based, on-site role with opportunities to mentor teams and influence SRE practices.

Qualifications

  • 7+ years of overall experience in Site Reliability Engineering / Production Support / Application Support.
  • Strong hands-on experience in Python scripting.
  • Proven experience in L3/L4 production support.
  • Strong knowledge of Linux/Unix environments.
  • Strong hands-on experience with SQL and database troubleshooting.
  • Experience in incident management and major/high-impact incident handling.
  • Strong experience in RCA, problem management, and service restoration.
  • Experience with production monitoring, troubleshooting, log analysis, and operational tools.
  • Experience in automation and scripting to reduce manual operational activities.
  • Proven experience in handling/leading support teams and coordinating multiple technology teams during critical incidents.
  • Strong understanding of SLA, incident severity, escalation, change management, and IT service management processes.
  • Excellent communication and stakeholder management skills.
  • Ability to work under pressure and take ownership during critical production incidents.

Responsibilities

  • Drive and coordinate L3/L4 production support activities for critical business applications and services.
  • Own the end-to-end incident management and service restoration process for high-impact incidents.
  • Assess incident severity, priority, business impact, customer impact, and risk and ensure appropriate escalation.
  • Lead troubleshooting and restoration activities for Sev1Sev4 incidents.
  • Coordinate with application, infrastructure, database, network, cloud, and other technology teams to restore services within agreed SLAs.
  • Perform detailed root‑cause analysis (RCA) and drive corrective and preventive actions.
  • Utilize Python scripting for automation, troubleshooting, monitoring, operational improvements, and repetitive task reduction.
  • Develop and maintain scripts/tools to automate incident resolution, health checks, operational activities, and service recovery processes.
  • Troubleshoot production issues using Linux/Unix commands, SQL queries, logs, monitoring tools, and system/application data.
  • Investigate and remediate customer/client data issues and coordinate with relevant technology teams for resolution.
  • Perform activities such as batch restarts, service restarts, routing changes, contingency procedures, and controlled recovery actions as required.
  • Engage with external software/hardware vendors when specialized technical support is required.
  • Drive High Impact Incident Communications, providing timely updates on incident status, business/customer impact, troubleshooting progress, and service restoration.
  • Ensure incident tickets contain accurate and complete information, including impact, timeline, actions taken, participants, resolution, and RCA details.
  • Work closely with Problem Management teams to provide detailed incident information and support creation of Problem Records.
  • Identify and document known errors, repeatable incidents, and recurring production issues in the Known Error Database.
  • Conduct regular incident reviews to identify trends, recurring issues, and opportunities for service improvement.
  • Create and maintain standard operating procedures, technical documentation, troubleshooting guides, and incident playbooks.
  • Identify opportunities for task automation, tooling improvements, and operational efficiency.
  • Participate in resiliency exercises and Chaos Engineering activities to improve application and infrastructure reliability.
  • Support audit, compliance, and risk remediation activities related to production operations.
  • Participate in high-risk and complex technology changes, including Permit‑to‑Operate / change governance activities.
  • Monitor service reliability and proactively identify potential production risks and failure points.
  • Ensure appropriate access and temporary access requests are reviewed and managed in accordance with defined procedures.
  • Lead or coordinate support teams during critical incidents and ensure effective collaboration across technology functions.
  • Mentor and guide team members on incident management, troubleshooting, production support, automation, and SRE best practices.
  • Ensure the support team follows defined processes, SLAs, escalation procedures, and operational standards.
  • Drive continuous improvement initiatives to increase service availability, reliability, automation, and operational efficiency.

Skills

Python scripting
Linux/Unix
SQL
Incident management
Root cause analysis
Automation
SRE/DevOps
Cloud platforms (AWS/Azure/GCP)
Communication
Team leadership
SLAs / escalation
Problem management

Tools

Monitoring tools
Log analysis tools
Operational tools

Job description

SRE Lead

Exp

7-12 Yrs


Location

Bangalore


Job Description

We are looking for an experienced Site Reliability Engineer (SRE) with strong expertise in Python scripting, Linux/Unix, SQL, production support, incident management, and automation. The ideal candidate should have hands‑on experience working in L3/L4 support environments, managing critical production incidents, driving root‑cause analysis, and leading/handling support teams.


Key Responsibilities


  • Drive and coordinate L3/L4 production support activities for critical business applications and services.

  • Own the end-to-end incident management and service restoration process for high-impact incidents.

  • Assess incident severity, priority, business impact, customer impact, and risk and ensure appropriate escalation.

  • Lead troubleshooting and restoration activities for Sev1Sev4 incidents.

  • Coordinate with application, infrastructure, database, network, cloud, and other technology teams to restore services within agreed SLAs.

  • Perform detailed root‑cause analysis (RCA) and drive corrective and preventive actions.

  • Utilize Python scripting for automation, troubleshooting, monitoring, operational improvements, and repetitive task reduction.

  • Develop and maintain scripts/tools to automate incident resolution, health checks, operational activities, and service recovery processes.

  • Troubleshoot production issues using Linux/Unix commands, SQL queries, logs, monitoring tools, and system/application data.

  • Investigate and remediate customer/client data issues and coordinate with relevant technology teams for resolution.

  • Perform activities such as batch restarts, service restarts, routing changes, contingency procedures, and controlled recovery actions as required.

  • Engage with external software/hardware vendors when specialized technical support is required.

  • Drive High Impact Incident Communications, providing timely updates on incident status, business/customer impact, troubleshooting progress, and service restoration.

  • Ensure incident tickets contain accurate and complete information, including impact, timeline, actions taken, participants, resolution, and RCA details.

  • Work closely with Problem Management teams to provide detailed incident information and support creation of Problem Records.

  • Identify and document known errors, repeatable incidents, and recurring production issues in the Known Error Database.

  • Conduct regular incident reviews to identify trends, recurring issues, and opportunities for service improvement.

  • Create and maintain standard operating procedures, technical documentation, troubleshooting guides, and incident playbooks.

  • Identify opportunities for task automation, tooling improvements, and operational efficiency.

  • Participate in resiliency exercises and Chaos Engineering activities to improve application and infrastructure reliability.

  • Support audit, compliance, and risk remediation activities related to production operations.

  • Participate in high-risk and complex technology changes, including Permit‑to‑Operate / change governance activities.

  • Monitor service reliability and proactively identify potential production risks and failure points.

  • Ensure appropriate access and temporary access requests are reviewed and managed in accordance with defined procedures.

  • Lead or coordinate support teams during critical incidents and ensure effective collaboration across technology functions.

  • Mentor and guide team members on incident management, troubleshooting, production support, automation, and SRE best practices.

  • Ensure the support team follows defined processes, SLAs, escalation procedures, and operational standards.

  • Drive continuous improvement initiatives to increase service availability, reliability, automation, and operational efficiency.


Mandatory Skills


  • 7+ years of overall experience in Site Reliability Engineering / Production Support / Application Support.

  • Strong hands‑on experience in Python scripting.

  • Proven experience in L3/L4 production support.

  • Strong knowledge of Linux/Unix environments.

  • Strong hands‑on experience with SQL and database troubleshooting.

  • Experience in incident management and major/high-impact incident handling.

  • Strong experience in RCA, problem management, and service restoration.

  • Experience with production monitoring, troubleshooting, log analysis, and operational tools.

  • Experience in automation and scripting to reduce manual operational activities.

  • Proven experience in handling/leading support teams and coordinating multiple technology teams during critical incidents.

  • Strong understanding of SLA, incident severity, escalation, change management, and IT service management processes.

  • Excellent communication and stakeholder management skills.

  • Ability to work under pressure and take ownership during critical production incidents.


Good to Have


  • Experience with SRE/DevOps practices and tools.

  • Experience with cloud platforms such as AWS, Azure, or GCP.

  • Experience with monitoring and observability tools.

  • Knowledge of CI/CD and deployment processes.

  • Exposure to Chaos Engineering and resiliency testing.

  • ITIL / relevant industry certification.

  • Experience supporting large-scale, business-critical applications.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

3across • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

On-site
INR 2,500,000 - 4,000,000
SRE Developer
SRE Developer

Cloudxtreme • Hyderabad

On-site
INR 1,500,000 - 2,400,000
Site Reliability Engineering Lead (Application SRE Lead)
Site Reliability Engineering Lead (Application SRE Lead)

Hirexa Solutions • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Application SRE
Application SRE

Cloudxtreme • Pune District

On-site
INR 1,200,000 - 1,800,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
Site Reliability Engineering Lead_Truist
Site Reliability Engineering Lead_Truist

Infosys • Bengaluru

On-site
INR 4,000,000 - 7,000,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing • Pune District

On-site
INR 2,250,000 - 2,750,000