SRE / Production Engineering SRE, Production engineering,Kubernetes, Terraform, CI/CD pipelines, Observability (Prometheus/Grafana)
Preferred Qualifications
- Experience designing and implementing SLO/SLI frameworks, error budgets, and reliability KPIs across multiple teams.
- Strong background in observability practices (monitoring, alerting, logging, tracing) and building actionable operational dashboards.
- Expertise in release/change management practices that improve deployment safety and reduce production incidents.
- Experience leading cross-team reliability programs, influencing stakeholders, and driving measurable improvements in uptime and MTTR.
- Track record of mentoring engineers and setting engineering standards for operational excellence and automation at scale.
Key Responsibilities
Reliability & Production Ownership
- Lead production engineering practices to ensure high availability, scalability, and performance across services and platforms.
- Define and drive SLOs/SLIs, error budgets, capacity planning, and reliability roadmaps aligned to business priorities.
- Partner with engineering teams to design resilient architectures and reduce operational risk through proactive improvements.
Incident Management & Operational Excellence
- Own incident response processes (on-call readiness, triage, escalation, communication) and lead major incident bridges when needed.
- Drive blameless postmortems, root-cause analysis, and corrective/preventive actions to prevent recurrence.
- Establish operational runbooks, playbooks, and production readiness reviews for new releases and changes.
Cloud Operations & Automation
- Lead cloud operations to ensure secure, cost-effective, and reliable environments across regions/accounts/subscriptions.
- Identify toil and implement automation to improve deployment safety, recovery time, and operational efficiency.
- Standardize operational tooling and workflows to improve service health, change success rate, and MTTR.
Leadership & Collaboration
- Mentor engineers and influence cross-functional teams to adopt reliability engineering best practices.
- Provide technical leadership in prioritization, execution planning, and stakeholder communication for reliability initiatives.
Minimum Qualifications
- BTECH, MTECH, MCA, or MSC in Computer Science, IT, or a related field (or equivalent practical experience).
- 12–14 years of experience in SRE, Production Engineering, or Reliability Engineering roles supporting large-scale systems.
- Strong hands‑on experience in cloud operations, incident management, and production support for critical services.
- Proven ability to drive automation initiatives that reduce manual effort and improve system reliability.
- Demonstrated experience leading operational processes such as on‑call, postmortems, and production readiness practices.