Site Reliability Engineering Lead (Application SRE Lead)

Hirexa Solutions

Bengaluru

Hybrid

INR 3,500,000 - 7,000,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hirexa Solutions in Hyderabad is seeking a Site Reliability Engineering Lead (Application SRE Lead) to drive reliability for our critical applications. This role combines individual contributor work (80%) with people leadership (20%), requiring 8+ years of hands-on SRE, production engineering, or platform operations experience, and strong incident management skills.

You will lead incident bridges, design SLI/SLO frameworks, and collaborate across cloud platforms (GCP/Azure/AWS) and CI/CD

Qualifications

  • 8+ years of experience in Application Support, Site Reliability Engineering, Production Engineering, or Platform Operations.
  • Minimum 2-4 years of experience managing technical teams in large-scale production environments.
  • Hands-on experience supporting highly available, business-critical applications.
  • Strong understanding of SRE principles, Reliability Engineering, Operational Excellence, Error Budgets, SLIs, SLOs, Availability Engineering, Capacity Planning.

Responsibilities

  • Delivery & Technical Contribution (80%)
  • Reliability Engineering: Ensure overall platform reliability, availability, and performance.
  • Production Support: Lead Sev1 and Sev2 incident investigations; drive service restoration.
  • Application Operations: Troubleshoot complex issues; monitor transaction flows; ensure authentication services.
  • Deployments & Release Management: Support deployments; ensure rollout and rollback; governance.

Skills

SRE principles
Reliability Engineering
Operational Excellence
Error Budgets
SLIs
SLOs
Kubernetes
Docker
Service Mesh
Incident Management

Tools

OAuth 2.0
OIDC
SAML
MFA
Okta
Transmit Security
IAM
ServiceNow
Remedy
JIRA

Job description

Details
Job Title*

Site Reliability Engineering Lead (Application SRE Lead)

Role Type*

Individual Contributor (80%) + People Leadership (20%)

Experience Level*

Min: 8 Years

Location*

Hyderabad

Work Type*

Full-time

Shift Requirement*

24x7 Production Support and On-call Rotation

Coding Requirement*

Mandatory

Technical Requirements
SRE & Production Engineering
  • 8+ years of experience in Application Support, Site Reliability Engineering, Production Engineering, or Platform Operations.
  • Minimum 2-4 years of experience managing technical teams in large-scale production environments.
  • Hands-on experience supporting highly available, business-critical applications.
  • Strong understanding of:
    • SRE principles
    • Reliability Engineering
    • Operational Excellence
    • Error Budgets
    • Service Level Indicators (SLIs)
    • Service Level Objectives (SLOs)
    • Availability Engineering
    • Capacity Planning
  • Expertise managing production incidents, service outages, and major incident bridges.
Authentication & Security Platforms

Strong understanding and operational support experience in:

  • OAuth 2.0
  • OpenID Connect (OIDC)
  • SAML
  • Multi-Factor Authentication (MFA)
  • Session Management
  • Token Lifecycle Management
  • API Authentication
  • Okta
  • Transmit Security
  • Identity and Access Management (IAM)
Application Troubleshooting

Hands-on troubleshooting expertise involving:

  • Java applications
  • .NET Core applications
  • REST APIs
  • Microservices
  • Batch processing applications
  • Application latency issues
  • Memory leaks
  • Thread contention
  • Authentication failures
  • Service-to-service communication issues
Cloud Platforms

Experience managing applications hosted on:

  • Google Cloud Platform (Preferred)
  • Microsoft Azure
  • Amazon Web Services
  • PCF (Pivotal Cloud Foundry)

Areas of expertise:

  • Application hosting
  • Networking
  • Cloud security
  • Service reliability
  • Monitoring
  • Disaster recovery
Containers & Microservices

Strong experience with:

  • Kubernetes
  • Docker
  • Helm
  • Service Mesh concepts
  • Container troubleshooting
  • Pod lifecycle management
  • Ingress controllers
  • Service discovery
  • Scaling strategies
Observability & Monitoring

Hands-on expertise in:

  • Splunk
  • AppDynamics
  • Datadog
  • Grafana
  • ThousandEyes
  • ITRS Geneos
  • MoogSoft
  • AppMetrics

Experience with:

  • Log Analytics
  • Distributed Tracing
  • Metrics Monitoring
  • Synthetic Monitoring
  • Alert Tuning
  • Dashboard Creation
  • Event Correlation
Database & Messaging

Strong understanding of:

  • MongoDB
  • PostgreSQL
  • SQL Query Optimization
  • Database Performance Tuning
  • Replication Issues
  • Database Connectivity Troubleshooting
  • Kafka Monitoring and Troubleshooting
Automation & Scripting

Strong coding and automation skills using:

  • Shell Scripting
  • Python
  • PowerShell
  • Go (Preferred)
  • Java (Preferred)

Experience automating:

  • Operational runbooks
  • Deployment validation
  • Monitoring
  • Incident remediation
  • Service recovery procedures
CI/CD & Release Management

Experience with:

  • Harness
  • Bamboo
  • Bitbucket Pipelines
  • GitHub Actions
  • Jenkins
  • Azure DevOps

Strong understanding of:

  • CI/CD
  • Release Engineering
  • Deployment Strategies
  • Blue-Green Deployments
  • Canary Deployments
  • Rollback Procedures
ITIL & ITSM

Strong experience in:

  • Incident Management
  • Problem Management
  • Change Management
  • Release Management
  • Knowledge Management

Tools:

  • ServiceNow
  • Remedy
  • JIRA
Key Responsibilities
Delivery & Technical Contribution (80%)
Reliability Engineering
  • Ensure overall platform reliability, availability, and performance.
  • Drive continuous improvements to reduce incidents and operational risks.
  • Design and implement SLI/SLO frameworks.
  • Monitor service health and proactively address burn-rate violations.
Production Support
  • Lead Sev1 and Sev2 incident investigations.
  • Drive service restoration and stakeholder communication.
  • Manage major incident bridges and technical war rooms.
  • Perform Root Cause Analysis and post-mortem reviews.
Application Operations
  • Troubleshoot complex application, infrastructure, database, and network issues.
  • Support authentication and authorization services globally.
  • Monitor business-critical transaction flows.
  • Build advanced monitoring dashboards and synthetic health checks.
Deployments & Release Management
  • Support production deployments and releases.
  • Lead deployment readiness assessments.
  • Ensure successful rollout and rollback execution.
  • Drive release governance processes.
Automation & Engineering Excellence
  • Automate operational processes and runbooks.
  • Implement self-healing and auto-remediation capabilities.
  • Identify toil reduction opportunities.
  • Improve MTTD and MTTR metrics.
Agile Delivery
  • Participate in sprint planning.
  • Own and deliver assigned epics and user stories.
  • Review technical solutions and implementation approaches.
  • Ensure operational readiness for new platform capabilities.
Leadership & Team Management (20%)
Team Leadership
  • Lead, mentor, and coach SRE engineers.
  • Conduct technical reviews and guidance sessions.
  • Support career development initiatives.
  • Build a culture of operational excellence.
Delivery Governance
  • Ensure SLA, SLO, and operational commitments are consistently achieved.
  • Monitor service delivery metrics.
  • Review team performance and workload distribution.
  • Drive capacity and resource planning.
Stakeholder Management
  • Act as primary escalation point for critical incidents.
  • Communicate effectively with business, engineering, and executive leadership.
  • Manage client expectations during outages and major events.
Continuous Improvement
  • Drive service improvement initiatives.
  • Lead automation programs.
  • Improve reliability maturity across application portfolios.
  • Contribute to organizational SRE best practices.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Support Lead
Production Support Lead

Cloudxtreme • Hyderabad

On-site
INR 3,200,000 - 6,000,000
Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Hyderabad

On-site
INR 1,400,000 - 2,200,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Application SRE
Application SRE

Cloudxtreme • Pune District

On-site
INR 1,200,000 - 1,800,000
Site Reliability Engineering Lead_Truist
Site Reliability Engineering Lead_Truist

Infosys • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Bengaluru

On-site
INR 900,000 - 1,500,000
SRE Expert
SRE Expert

HCLTech • Bengaluru

On-site
INR 1,500,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

Nexcess • India

On-site
INR 2,000,000 - 4,200,000