Production Support Lead

Cloudxtreme

Hyderabad

On-site

INR 3,200,000 - 6,000,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cloudxtreme in Hyderabad is seeking a Senior SRE/Production Engineering lead to own platform reliability and availability. You will design SLI/SLO frameworks, monitor health, and drive improvements across cloud and on‑prem services.

You will coordinate Sev1/ Sev2 incidents, lead post‑mortems, and mentor a globally distributed team while delivering robust CI/CD and deployment governance.

Qualifications

  • Extensive experience in SRE/Production Engineering in large-scale environments.
  • Proven ability to manage technical teams and ensure reliability.
  • Hands-on troubleshooting of high-availability applications.

Responsibilities

  • Delivery & Technical Contribution (80%): ensure platform reliability, define and enforce SLI/SLO frameworks, monitor health, reduce incidents.
  • Production Support: lead Sev1/Sev2 investigations, coordinate war rooms and post-mortems.
  • Application Operations: troubleshoot complex issues across apps, infra, DB, and network.
  • Deployments & Release Management: support deployments, drive readiness and governance.
  • Automation & Engineering Excellence: automate runbooks, implement self-healing, reduce toil.
  • Agile Delivery: participate in planning, deliver epics and stories, ensure operational readiness.
  • Leadership & Team Management (20%): mentor SRE engineers, manage workload, and drive capacity planning.

Skills

SRE principles
Reliability Engineering
Operational Excellence
Error Budgets
SLIs
SLOs
Availability Engineering
Capacity Planning
Incident management
Team leadership

Tools

Kubernetes
Docker
Helm
Service Mesh
Ingress controllers
Service discovery
Monitoring
Splunk
Datadog
Grafana

Job description

Job Requirements
Technical Requirements
SRE & Production Engineering
  • 8+ years of experience in Application Support, Site Reliability Engineering, Production Engineering, or Platform Operations.
  • Minimum 2-4 years of experience managing technical teams in large-scale production environments.
  • Hands-on experience supporting highly available, business-critical applications.
  • Strong understanding of:
    • SRE principles
    • Reliability Engineering
    • Operational Excellence
    • Error Budgets
    • Service Level Indicators (SLIs)
    • Service Level Objectives (SLOs)
    • Availability Engineering
    • Capacity Planning
  • Expertise managing production incidents, service outages, and major incident bridges.
Authentication & Security Platforms

Strong understanding and operational support experience in:

  • OAuth 2.0
  • OpenID Connect (OIDC)
  • SAML
  • Multi-Factor Authentication (MFA)
  • Session Management
  • Token Lifecycle Management
  • API Authentication
  • Okta
  • Transmit Security
  • Identity and Access Management (IAM)
Application Troubleshooting

Hands-on troubleshooting expertise involving:

  • Java applications
  • .NET Core applications
  • REST APIs
  • Microservices
  • Batch processing applications
  • Application latency issues
  • Memory leaks
  • Thread contention
  • Authentication failures
  • Service-to-service communication issues
Cloud Platforms

Experience managing applications hosted on:

  • Google Cloud Platform (Preferred)
  • Microsoft Azure
  • Amazon Web Services
  • PCF (Pivotal Cloud Foundry)

Areas of expertise:

  • Application hosting
  • Networking
  • Cloud security
  • Service reliability
  • Monitoring
  • Disaster recovery
Containers & Microservices

Strong experience with:

  • Kubernetes
  • Docker
  • Helm
  • Service Mesh concepts
  • Container troubleshooting
  • Pod lifecycle management
  • Ingress controllers
  • Service discovery
  • Scaling strategies
Observability & Monitoring

Hands-on expertise in:

  • Splunk
  • AppDynamics
  • Datadog
  • Grafana
  • ThousandEyes
  • ITRS Geneos
  • MoogSoft
  • AppMetrics

Experience with:

  • Log Analytics
  • Distributed Tracing
  • Metrics Monitoring
  • Synthetic Monitoring
  • Alert Tuning
  • Dashboard Creation
  • Event Correlation
Database & Messaging

Strong understanding of:

  • MongoDB
  • PostgreSQL
  • SQL Query Optimization
  • Database Performance Tuning
  • Replication Issues
  • Database Connectivity Troubleshooting
  • Kafka Monitoring and Troubleshooting
Automation & Scripting

Strong coding and automation skills using:

  • Shell Scripting
  • Python
  • PowerShell
  • Go (Preferred)
  • Java (Preferred)

Experience automating:

  • Operational runbooks
  • Deployment validation
  • Monitoring
  • Incident remediation
  • Service recovery procedures
CI/CD & Release Management

Experience with:

  • Harness
  • Bamboo
  • Bitbucket Pipelines
  • GitHub Actions
  • Jenkins
  • Azure DevOps

Strong understanding of:

  • CI/CD
  • Release Engineering
  • Deployment Strategies
  • Blue-Green Deployments
  • Canary Deployments
  • Rollback Procedures
ITIL & ITSM

Strong experience in:

  • Incident Management
  • Problem Management
  • Change Management
  • Release Management
  • Knowledge Management

Tools:

  • ServiceNow
  • Remedy
  • JIRA
Key Responsibilities
Delivery & Technical Contribution (80%)
Reliability Engineering
  • Ensure overall platform reliability, availability, and performance.
  • Drive continuous improvements to reduce incidents and operational risks.
  • Design and implement SLI/SLO frameworks.
  • Monitor service health and proactively address burn-rate violations.
Production Support
  • Lead Sev1 and Sev2 incident investigations.
  • Drive service restoration and stakeholder communication.
  • Manage major incident bridges and technical war rooms.
  • Perform Root Cause Analysis and post-mortem reviews.
Application Operations
  • Troubleshoot complex application, infrastructure, database, and network issues.
  • Support authentication and authorization services globally.
  • Monitor business-critical transaction flows.
  • Build advanced monitoring dashboards and synthetic health checks.
Deployments & Release Management
  • Support production deployments and releases.
  • Lead deployment readiness assessments.
  • Ensure successful rollout and rollback execution.
  • Drive release governance processes.
Automation & Engineering Excellence
  • Automate operational processes and runbooks.
  • Implement self-healing and auto-remediation capabilities.
  • Identify toil reduction opportunities.
  • Improve MTTD and MTTR metrics.
Agile Delivery
  • Participate in sprint planning.
  • Own and deliver assigned epics and user stories.
  • Review technical solutions and implementation approaches.
  • Ensure operational readiness for new platform capabilities.
Leadership & Team Management (20%)
Team Leadership
  • Lead, mentor, and coach SRE engineers.
  • Conduct technical reviews and guidance sessions.
  • Support career development initiatives.
  • Build a culture of operational excellence.
Delivery Governance
  • Ensure SLA, SLO, and operational commitments are consistently achieved.
  • Monitor service delivery metrics.
  • Review team performance and workload distribution.
  • Drive capacity and resource planning.
Stakeholder Management
  • Act as primary escalation point for critical incidents.
  • Communicate effectively with business, engineering, and executive leadership.
  • Manage client expectations during outages and major events.
Continuous Improvement
  • Drive service improvement initiatives.
  • Lead automation programs.
  • Improve reliability maturity across application portfolios.
  • Contribute to organizational SRE best practices.
Soft Skills
  • Excellent communication and stakeholder management
  • Strong leadership and mentoring capabilities
  • High ownership and accountability
  • Strategic problem-solving mindset
  • Ability to make decisions under pressure
  • Customer-centric attitude
  • Conflict resolution and collaboration skills
  • Strong documentation practices
  • Ability to lead globally distributed teams
Readiness & Work Conditions
  • Ready to work from office 5 days a week.
  • Ready to support 24x7 Production Support operations.
  • Comfortable working in rotational shifts and on-call support.
  • Ready to support mission‑critical customer‑facing platforms.
  • Ready to upskill continuously on emerging technologies and cloud platforms.
Preferred Certifications
Cloud
  • Google Cloud Associate Cloud Engineer (Preferred)
  • Google Professional Cloud Architect
  • Microsoft Azure Administrator
  • AWS Solutions Architect Associate
Reliability & Operations
  • ITIL Foundation
  • Certified Kubernetes Administrator (CKA)
  • Splunk Certified Power User/Admin
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
Site Reliability Engineer (SRE) – GCP Platform
Site Reliability Engineer (SRE) – GCP Platform

ITC Infotech • Bengaluru

On-site
INR 900,000 - 1,300,000
SRE- Production Support
SRE- Production Support

Cloudxtreme • Hyderabad

On-site
INR 2,400,000 - 3,600,000
Site Reliability Engineering (SRE) Lead
Site Reliability Engineering (SRE) Lead

SID Global Solutions • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Senior Manager System Reliability Engineering
Senior Manager System Reliability Engineering

GE Vernova, Inc. • Hyderabad

On-site
INR 4,000,000 - 6,000,000
Relocation Assistance Provided: Yes
Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000
SRE Lead
SRE Lead

Acldigital • Ahmedabad District

On-site
INR 1,500,000 - 2,000,000
Technical Support Engineer/SRE
Technical Support Engineer/SRE

Boldtek • India

On-site
INR 900,000 - 1,500,000
Hybrid work model
Growth opportunities
SRE
SRE

Jobtailor • Chennai District

On-site
INR 2,500,000 - 4,500,000