Job Requirements
Technical Requirements
SRE & Production Engineering
- 8+ years of experience in Application Support, Site Reliability Engineering, Production Engineering, or Platform Operations.
- Minimum 2-4 years of experience managing technical teams in large-scale production environments.
- Hands-on experience supporting highly available, business-critical applications.
- Strong understanding of:
- SRE principles
- Reliability Engineering
- Operational Excellence
- Error Budgets
- Service Level Indicators (SLIs)
- Service Level Objectives (SLOs)
- Availability Engineering
- Capacity Planning
- Expertise managing production incidents, service outages, and major incident bridges.
Authentication & Security Platforms
Strong understanding and operational support experience in:
- OAuth 2.0
- OpenID Connect (OIDC)
- SAML
- Multi-Factor Authentication (MFA)
- Session Management
- Token Lifecycle Management
- API Authentication
- Okta
- Transmit Security
- Identity and Access Management (IAM)
Application Troubleshooting
Hands-on troubleshooting expertise involving:
- Java applications
- .NET Core applications
- REST APIs
- Microservices
- Batch processing applications
- Application latency issues
- Memory leaks
- Thread contention
- Authentication failures
- Service-to-service communication issues
Cloud Platforms
Experience managing applications hosted on:
- Google Cloud Platform (Preferred)
- Microsoft Azure
- Amazon Web Services
- PCF (Pivotal Cloud Foundry)
Areas of expertise:
- Application hosting
- Networking
- Cloud security
- Service reliability
- Monitoring
- Disaster recovery
Containers & Microservices
Strong experience with:
- Kubernetes
- Docker
- Helm
- Service Mesh concepts
- Container troubleshooting
- Pod lifecycle management
- Ingress controllers
- Service discovery
- Scaling strategies
Observability & Monitoring
Hands-on expertise in:
- Splunk
- AppDynamics
- Datadog
- Grafana
- ThousandEyes
- ITRS Geneos
- MoogSoft
- AppMetrics
Experience with:
- Log Analytics
- Distributed Tracing
- Metrics Monitoring
- Synthetic Monitoring
- Alert Tuning
- Dashboard Creation
- Event Correlation
Database & Messaging
Strong understanding of:
- MongoDB
- PostgreSQL
- SQL Query Optimization
- Database Performance Tuning
- Replication Issues
- Database Connectivity Troubleshooting
- Kafka Monitoring and Troubleshooting
Automation & Scripting
Strong coding and automation skills using:
- Shell Scripting
- Python
- PowerShell
- Go (Preferred)
- Java (Preferred)
Experience automating:
- Operational runbooks
- Deployment validation
- Monitoring
- Incident remediation
- Service recovery procedures
CI/CD & Release Management
Experience with:
- Harness
- Bamboo
- Bitbucket Pipelines
- GitHub Actions
- Jenkins
- Azure DevOps
Strong understanding of:
- CI/CD
- Release Engineering
- Deployment Strategies
- Blue-Green Deployments
- Canary Deployments
- Rollback Procedures
ITIL & ITSM
Strong experience in:
- Incident Management
- Problem Management
- Change Management
- Release Management
- Knowledge Management
Tools:
Key Responsibilities
Delivery & Technical Contribution (80%)
Reliability Engineering
- Ensure overall platform reliability, availability, and performance.
- Drive continuous improvements to reduce incidents and operational risks.
- Design and implement SLI/SLO frameworks.
- Monitor service health and proactively address burn-rate violations.
Production Support
- Lead Sev1 and Sev2 incident investigations.
- Drive service restoration and stakeholder communication.
- Manage major incident bridges and technical war rooms.
- Perform Root Cause Analysis and post-mortem reviews.
Application Operations
- Troubleshoot complex application, infrastructure, database, and network issues.
- Support authentication and authorization services globally.
- Monitor business-critical transaction flows.
- Build advanced monitoring dashboards and synthetic health checks.
Deployments & Release Management
- Support production deployments and releases.
- Lead deployment readiness assessments.
- Ensure successful rollout and rollback execution.
- Drive release governance processes.
Automation & Engineering Excellence
- Automate operational processes and runbooks.
- Implement self-healing and auto-remediation capabilities.
- Identify toil reduction opportunities.
- Improve MTTD and MTTR metrics.
Agile Delivery
- Participate in sprint planning.
- Own and deliver assigned epics and user stories.
- Review technical solutions and implementation approaches.
- Ensure operational readiness for new platform capabilities.
Leadership & Team Management (20%)
Team Leadership
- Lead, mentor, and coach SRE engineers.
- Conduct technical reviews and guidance sessions.
- Support career development initiatives.
- Build a culture of operational excellence.
Delivery Governance
- Ensure SLA, SLO, and operational commitments are consistently achieved.
- Monitor service delivery metrics.
- Review team performance and workload distribution.
- Drive capacity and resource planning.
Stakeholder Management
- Act as primary escalation point for critical incidents.
- Communicate effectively with business, engineering, and executive leadership.
- Manage client expectations during outages and major events.
Continuous Improvement
- Drive service improvement initiatives.
- Lead automation programs.
- Improve reliability maturity across application portfolios.
- Contribute to organizational SRE best practices.
Soft Skills
- Excellent communication and stakeholder management
- Strong leadership and mentoring capabilities
- High ownership and accountability
- Strategic problem-solving mindset
- Ability to make decisions under pressure
- Customer-centric attitude
- Conflict resolution and collaboration skills
- Strong documentation practices
- Ability to lead globally distributed teams
Readiness & Work Conditions
- Ready to work from office 5 days a week.
- Ready to support 24x7 Production Support operations.
- Comfortable working in rotational shifts and on-call support.
- Ready to support mission‑critical customer‑facing platforms.
- Ready to upskill continuously on emerging technologies and cloud platforms.
Preferred Certifications
Cloud
- Google Cloud Associate Cloud Engineer (Preferred)
- Google Professional Cloud Architect
- Microsoft Azure Administrator
- AWS Solutions Architect Associate
Reliability & Operations
- ITIL Foundation
- Certified Kubernetes Administrator (CKA)
- Splunk Certified Power User/Admin