- Strong experience in enterprise Disaster Recovery planning, governance, and program coordination.
- Good understanding of Business Continuity Management, BCP/DR lifecycle, IT service continuity, and resilience frameworks.
- Ability to define, track, and report DR KPIs such as RTO, RPO, test effectiveness, readiness score, audit gaps, and remediation status.
- Experience preparing DR dashboards, governance reports, executive summaries, and readiness/compliance packs.
Job Description
Must Have Technical/Functional Skills
Disaster Recovery Planning & Governance
- Strong experience in enterprise Disaster Recovery planning, governance, and program coordination.
- Good understanding of Business Continuity Management, BCP/DR lifecycle, IT service continuity, and resilience frameworks.
- Ability to define, track, and report DR KPIs such as RTO, RPO, test effectiveness, readiness score, audit gaps, and remediation status.
- Experience preparing DR dashboards, governance reports, executive summaries, and readiness/compliance packs.
DR Strategy, Architecture & Resilience
- Ability to develop and maintain enterprise DR frameworks, standards, operating models, and recovery methodologies.
- Experience conducting application, infrastructure, data flow, and dependency assessments across business-critical services.
- Good knowledge of DR solutions across on-premises, cloud, and hybrid environments including AWS, Azure, and GCP.
- Ability to recommend resilience improvements for critical business services and support multi-year DR roadmap initiatives.
Recovery Planning, Documentation & Runbooks
- Hands-on experience creating and maintaining DR plans, runbooks, SOPs, recovery procedures, escalation matrices, and communication templates.
- Ability to maintain application-to-infrastructure dependency mapping, DR asset inventory, and recovery ownership matrix.
- Experience coordinating periodic reviews and updates of DR documentation with application, infrastructure, security, and business teams.
DR Testing, Recovery Validation & Exercises
- Experience planning and coordinating tabletop exercises, simulation drills, partial failover testing, and full-scale recovery testing.
- Ability to validate RTO and RPO compliance during exercises and document test outcomes, risks, gaps, and remediation actions.
- Experience coordinating failover and failback activities with infrastructure, application, cloud, backup, network, and security teams.
Incident Response, Risk & Compliance
- Experience supporting DR activation during major incidents, disasters, cyber events, ransomware scenarios, regional disruptions, and infrastructure failures.
- Ability to coordinate command center activities, recovery milestones, stakeholder updates, and executive/customer communication during recovery events.
- Good understanding of IT risk assessments, audit evidence preparation, compliance reporting, and readiness assessment support.
- Familiarity with ISO 22301, SOC, HIPAA, PCI-DSS, ITIL practices, change management, incident management, and major incident recovery.
Required Technical Skills
- Disaster Recovery Planning & Governance; Business Continuity Management (BCP/DR); Enterprise Infrastructure Recovery.
- Cloud DR across AWS, Azure, and GCP; backup and recovery technologies; replication technologies; DR testing and validation.
- ITIL Service Continuity Management; risk assessment and compliance; recovery runbook development; infrastructure and application dependency mapping.
- Change management; executive reporting and governance; incident management and major incident recovery; regulatory compliance exposure.
Roles & Responsibilities
DR Governance & Program Management
- Lead and coordinate enterprise DR governance activities.
- Define and monitor DR KPIs including Recovery Time Objective (RTO), Recovery Point Objective (RPO), DR testing effectiveness, and compliance/audit metrics.
- Develop and maintain DR governance dashboards and executive reports.
- Drive compliance with organizational policies, regulatory requirements, and industry standards.
Disaster Recovery Strategy & Architecture
- Develop and maintain enterprise Disaster Recovery frameworks, standards, and methodologies.
- Support multi-year DR roadmap initiatives and technology modernization programs.
- Conduct infrastructure and application dependency assessments.
- Evaluate DR solutions across on-premises, cloud, and hybrid environments.
- Recommend resilience improvements for critical business services.
DR Planning & Documentation
- Develop and maintain DR plans for critical applications and infrastructure.
- Create and standardize runbooks, SOPs, recovery procedures, and escalation matrices.
- Coordinate periodic reviews and updates of DR documentation.
- Maintain application-to-infrastructure dependency mappings and DR asset inventories.
DR Testing & Recovery Exercises
- Plan and coordinate DR exercises including tabletop exercises, simulation drills, partial failover testing, and full-scale recovery testing.
- Validate RTO and RPO compliance during exercises.
- Lead post-test reviews and remediation tracking.
- Coordinate failover and failback activities with infrastructure and application teams.
Incident Response & Recovery Coordination
- Support DR activation during major incidents and disasters.
- Coordinate cross-functional recovery teams and command centre activities.
- Track recovery milestones and communicate status updates to key stakeholders.
- Provide executive and customer communication during recovery events.
Risk, Audit & Compliance
- Coordinate IT risk assessments related to DR and resiliency.
- Support internal, customer, and regulatory audits.
- Provide DR evidence, compliance reports, and readiness assessments.
- Track remediation of identified gaps and risks.
Vendor & Stakeholder Management
- Collaborate with cloud providers, data centre teams, infrastructure vendors, and third-party service providers.
- Coordinate vendor participation in DR testing and recovery events.
- Manage stakeholder communications across executive, operational, and technical teams.
Continuous Improvement & Training
- Drive lessons learned from incidents and DR exercises.
- Identify opportunities for improving resiliency and recovery capabilities.
- Conduct DR awareness and training sessions for technical and business teams.
- Expand recovery scenarios covering cyber incidents, ransomware, infrastructure failures, and regional disruptions.
Salary Range: $100,000 -$110,000 year
TCS Employee Benefits Summary
- Discretionary Annual Incentive.
- Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.
- Family Support: Maternal & Parental Leaves.
- Insurance Options: Auto & Home Insurance, Identity Theft Protection.
- Convenience & Professional Growth: Commuter Benefits & C ertification & amp; Training Reimbursement.
- Time Off: Vacation, Time Off, Sick Leave & Holidays.
- Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.
Qualifications:
BACHELOR OF COMPUTER SCIENCE