Job Description (JD) - IT Disaster Recovery and Availability Manager
Note: The candidate should be flexible and available to work on weekends for DR execution.
IT Disaster recovery:
- Participating in disaster Recovery Governance & Planning
- Maintain and periodically review the inventory of critical business and IT services used for Disaster Recovery planning and testing.
- Manage, annual Disaster Recovery and High Availability testing calendar.
- Ensure DR plans align with business continuity requirements and regulatory obligations.
- Ensure ITSCM plans, risks, controls, and activities remain aligned with business continuity requirements.
- Assess continuity risks and recommend mitigation strategies.
- Lead end-to-end Disaster Recovery and High Availability testing activities.
- Coordinate pre-test planning, stakeholder communications, execution workshops, and test rehearsals.
- Manage test windows, bridge calls.
- Monitor runbook execution and ensure successful completion of recovery activities.
- Coordinate evidence collection, issue tracking, risk logging, and post-test reviews.
- Prepare executive-level DR testing reports and recommendations
- Facilitate coordination between technical teams, business stakeholders, management, and external vendors.
- Participate in crisis management and business continuity response activities.
Availability Management:
- Execute and govern the Availability Management process in line with agreed SLAs, OLAs, and service commitments.
- Establish and maintain service availability baselines and performance targets.
- Conduct regular service availability reviews with stakeholders.
- Monitor availability, uptime, and downtime of all applications and infrastructure services utilizing approved monitoring tools.
- Validate availability calculations and ensure data integrity across reporting platforms.
- Generate daily, weekly, and executive-level availability reports.
- Analyze service outages, performance degradation, availability trends, and recurring failures.
- Identify risks affecting service availability and develop mitigation plans.
- Recommend service reliability and resilience improvements.
- Validate exclusion of approved downtime from SLA calculations.
- Conduct service review meetings and availability governance discussions.
- Define monitoring requirements, availability targets, reporting standards, and governance controls for new services.
- Identify opportunities to improve service availability and operational efficiency.
- Drive root cause trending and proactive improvement initiatives.
- Submit recommendations to CSI governance boards.
- Measure effectiveness of implemented improvements.
- Enhance availability analytics using Power BI, ServiceNow, Dynatrace, SiteScope, SCOM, SolarWinds, or equivalent monitoring tools.
- Reduce manual effort through automation of KPI reporting and trend analysis.
Required Skill:
- Disaster Recovery Management
- Business Continuity Management (BCM)
- Business Impact Analysis (BIA)
- Availability Management
- RTO/RPO Management
- DR Runbook Management
- Power BI, ServiceNow (User level)
- Dynatrace, SiteScope, SolarWinds, (User level)
- Executive Reporting & Stakeholder Management.
KPI:
- DR and HA testing completion rate.
- DR testing success rate.
- Achievement of agreed RTO and RPO targets.
- Closure rate of DR findings and lessons learned.
- DR readiness compliance score.
- IT service availability percentage.
- SLA compliance achievement.
- Completion of availability improvement initiatives.
- Automation and reporting efficiency improvements