Site Reliability Engineer

Tribute Technology Career Center

Hyderabad

On-site

INR 900,000 - 1,500,000

Full time

12 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Tribute Technology is seeking a Site Reliability Engineer to join our Cloud Operations team in Hyderabad. You will participate in a 24x7 on-call rotation, monitor platform health with Datadog, and drive incident response and root-cause analysis with cross-functional teams.

You will contribute to automation, reliability improvements, and disaster recovery planning while documenting procedures and contributing to operational excellence across cloud platforms.

Qualifications

  • 2-4 years of experience in Site Reliability Engineering or Cloud Operations.
  • Experience responding to production incidents and building reliable systems.
  • Familiarity with observability and alerting best practices.

Responsibilities

  • Participate in a 24x7 on-call rotation as a primary responder for production incidents.
  • Triage, investigate, and mitigate alerts using runbooks and escalation procedures.
  • Coordinate with engineering teams during major incident investigations.
  • Create and track incident timelines and actions in Jira.
  • Develop automation to reduce toil and improve reliability.
  • Support disaster recovery, capacity planning, and post-change monitoring.
  • Maintain documentation in Confluence and contribute to knowledge sharing.

Skills

Incident response
On-call rotation
Monitoring
Observability
Automation
Cloud Operations
AWS

Tools

Datadog
Jira
Confluence
AWS

Job description

ABOUT TRIBUTE TECHNOLOGY:

At Tribute Technology, we make end-of-life celebrations memorable, meaningful, and effortless through thoughtful and innovative technology solutions. Our mission is to help communities around the world celebrate life and pay tribute to those we love. Our comprehensive platform brings together software and technology to provide a fully integrated experience for all users, whether that is a family, a funeral home, or an online publisher. We are the market leader in the US and Canada, with global expansion plans and a growing international team of more than 400 individuals in the US, Canada, Philippines, Ukraine and India.

ABOUT YOU:

We are looking for a passionate and driven Site Reliability Engineer (SRE) with 2-4 years of experience to join our Cloud Operations team. This role is ideal for someone who enjoys solving operational challenges, responding to production incidents, improving system reliability, and building scalable monitoring and automation solutions.

As a member of the SRE team, you will play a critical role in maintaining the availability, performance, and resilience of Tribute's cloud platforms. You will participate in a 24x7 on-call rotation, investigate production incidents, contribute to root cause analysis, improve observability, and partner closely with Engineering, DevOps, Database, and Cloud Operations teams.

This role offers an excellent opportunity to develop expertise in Site Reliability Engineering, Incident Management, Cloud Operations, and Reliability Engineering practices.

ESSENTIAL DUTIES AND RESPONSIBILITIES:
Incident Management & On-Call Support:
  • Participate in a 24x7 on-call rotation as a primary responder for production incidents.
  • Triage, acknowledge, investigate, and mitigate alerts generated through monitoring platforms.
  • Respond to incidents using established runbooks, playbooks, and escalation procedures.
  • Escalate Priority 1 and Priority 2 incidents according to the Incident Management process.
  • Coordinate with engineering teams during major incident investigations.
  • Communicate incident status updates clearly and professionally to stakeholders.
  • Maintain detailed incident timelines and documentation throughout the incident lifecycle.
  • Support incident bridge coordination and follow established incident response processes.
Problem Management & Root Cause Analysis:
  • Participate in Post Incident Reviews (PIRs) and Root Cause Analysis (RCA) activities.
  • Document root causes, contributing factors, impact, timeline, and remediation actions.
  • Create and track follow-up action items through Jira.
  • Drive assigned corrective and preventive actions to completion.
  • Identify recurring incidents and reliability risks and propose long-term solutions.
  • Contribute to continuous improvement initiatives based on incident learnings.
Monitoring, Observability & Alert Management:
  • Monitor platform health through Datadog dashboards, alerts, logs, and metrics.
  • Maintain awareness of service health, availability, and performance indicators.
  • Improve alert quality by reducing noise and eliminating false positives.
  • Identify monitoring gaps and implement enhanced observability solutions.
  • Support Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational reporting.
  • Collaborate with engineering teams to establish meaningful health checks and alerting strategies.
Reliability Engineering & Automation:
  • Identify repetitive manual activities and operational toil.
  • Develop or contribute to automation solutions using scripting and tooling.
  • Assist with production readiness reviews for new applications and services.
  • Support disaster recovery, failover testing, and resilience initiatives.
  • Maintain and improve operational runbooks and standard operating procedures.
  • Recommend platform improvements that enhance reliability, scalability, and performance.
Cloud Operations Support:
  • Perform incident investigation and troubleshooting within AWS environments.
  • Support cloud infrastructure, application, networking, and platform-related investigations.
  • Participate in deployment validation and post-change monitoring activities.
  • Collaborate with DevOps and Engineering teams during production releases.
  • Assist in capacity planning and platform health assessments.
Operational Excellence & Documentation:
  • Maintain accurate documentation in Confluence.
  • Ensure incident records, Jira tickets, and action items remain updated.
  • Participate in shift handovers and knowledge-sharing activities.
  • Contribute to operational process improvements and team best practices.
  • Support knowledge management and training initiatives within the team.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Hirebridge LLC • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Lead Engineer – Site Reliability Engineering
Lead Engineer – Site Reliability Engineering

CBTS • Chennai District

On-site
INR 5,000,000 - 7,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Bahwan CyberTek • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Lead Engineer – Site Reliability Engineering
Lead Engineer – Site Reliability Engineering

cbtsindia • Chennai District

On-site
INR 1,800,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MontyCloud • Bengaluru

On-site
INR 2,500,000 - 5,000,000
Lead SRE
Lead SRE

Cvent • Gurugram District

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer Cloud Ops Hyderabad
Senior Site Reliability Engineer Cloud Ops Hyderabad

Seismic • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

Saika Technologies Inc. • Hyderabad, Bengaluru

Hybrid
INR 3,000,000 - 4,200,000