DXC Technology تعلن عن وظيفة مدير التعافي من الكوارث وتوفر تقنية المعلومات في الرياض
تفاصيل الوظيفة
شركة DXC Technology تعلن عن حاجة لمدير تخطيط استمرارية الأعمال وتعافي الكوارث وتوافر الخدمات (IT Disaster Recovery and Availability Manager) في الرياض، المملكة العربية السعودية.
المهام والمسؤوليات
- المشاركة في حوكمة وتخطيط التعافي من الكوارث.
- الاحتفاظ بقائمة الخدمات الحرجة للتعافي ومراجعتها دورياً.
- إدارة التقويم السنوي لاختبار التعافي والتوفر العالي.
- ضمان توافق خطط التعافي مع متطلبات استمرارية الأعمال واللوائح التنظيمية.
- ضمان توافق خطط وإجراءات إدارة استمرارية الخدمات التقنية (ITSCM) مع متطلبات استمرارية الأعمال.
- تقييم مخاطر الاستمرارية وتقديم توصيات للتخفيف منها.
- قيادة أنشطة اختبار التعافي من الكوارث والتوفر العالي من البداية حتى النهاية.
- التنسيق للتخطيط المسبق للاختبار والتواصل مع أصحاب المصلحة وجلسات العمل والتدريبات.
- إدارة فترات الاختبار والاتصالات عبر الجسور.
- مراقبة تنفيذ دليل التشغيل وضمان إكمال أنشطة التعافي بنجاح.
- التنسيق لجمع الأدلة وتتبع المشكلات وتسجيل المخاطر ومراجعات ما بعد الاختبار.
- إعداد تقارير اختبار التعافي على المستوى التنفيذي مع التوصيات.
- تسهيل التنسيق بين الفرق الفنية وأصحاب المصلحة والإدارة والبائعين الخارجيين.
- المشاركة في أنشطة إدارة الأزمات والاستجابة لاستمرارية الأعمال.
- تنفيذ وإدارة عملية إدارة التوافر وفقاً لاتفاقيات مستوى الخدمة (SLA) والالتزامات.
- إنشاء والحفاظ على خطوط الأساس لتوافر الخدمة وأهداف الأداء.
- إجراء مراجعات دورية لتوافر الخدمة مع أصحاب المصلحة.
- مراقبة توافر وارتفاع وانخفاض التطبيقات والبنية التحتية باستخدام أدوات المراقبة المعتمدة.
- التحقق من حسابات التوافر وضمان سلامة البيانات عبر منصات التقارير.
- إعداد تقارير التوافر اليومية والأسبوعية والتنفيذية.
- تحليل توقف الخدمة وتدهور الأداء واتجاهات التوافر والإخفاقات المتكررة.
- تحديد المخاطر المؤثرة على توافر الخدمة ووضع خطط التخفيف.
- تقديم توصيات لتحسين موثوقية الخدمة ومرونتها.
- التحقق من استبعاد فترة التوقف المعتمدة من حسابات مستوى الخدمة (SLA).
- عقد اجتماعات مراجعة الخدمة ومناقشات حوكمة التوافر.
- تحديد متطلبات المراقبة وأهداف التوافر ومعايير التقارير والضوابط الحوكمة للخدمات الجديدة.
- تحديد فرص تحسين توافر الخدمة وكفاءة التشغيل.
- دفع تحليل الجذور والمبادرات التحسينية الاستباقية.
- تقديم التوصيات لمجالس حوكمة التحسين المستمر (CSI).
- قياس فعالية التحسينات المطبقة.
- تحسين تحليلات التوافر باستخدام Power BI وServiceNow وDynatrace وSiteScope وSCOM وSolarWinds أو أدوات مشابهة.
- تقليل الجهد اليدوي من خلال أتمتة تقارير مؤشرات الأداء الرئيسية (KPI) وتحليل الاتجاهات.
الشروط والمتطلبات
- يجب أن يكون المرشح مرناً ومتاحاً للعمل في عطلات نهاية الأسبوع لتنفيذ التعافي من الكوارث.
المهارات المطلوبة
- إدارة التعافي من الكوارث (Disaster Recovery Management).
- إدارة استمرارية الأعمال (Business Continuity Management - BCM).
- تحليل تأثير الأعمال (Business Impact Analysis - BIA).
- إدارة التوافر (Availability Management).
- إدارة أهداف زمن الاسترداد / أهداف نقطة الاسترداد (RTO/RPO Management).
- إدارة دليل تشغيل التعافي (DR Runbook Management).
- Power BI، ServiceNow (مستوى مستخدم).
- Dynatrace، SiteScope، SolarWinds (مستوى مستخدم).
- التقارير التنفيذية وإدارة أصحاب المصلحة (Executive Reporting & Stakeholder Management).
- معدل إتمام اختبارات التعافي والتوفر العالي.
- معدل نجاح اختبار التعافي.
- تحقيق أهداف RTO و RPO المتفق عليها.
- معدل إغلاق نتائج التعافي والدروس المستفادة.
- درجة الامتثال الجاهزية للتعافي.
- نسبة توافر خدمات تقنية المعلومات.
- تحقيق الامتثال لاتفاقيات مستوى الخدمة (SLA).
- إتمام مبادرات تحسين التوافر.
- تحسينات كفاءة الأتمتة والتقارير.
عرض النص الأصلي للإعلان
Job Description (JD) - IT Disaster Recovery and Availability Manager
Note: The candidate should be flexible and available to work on weekends for DR execution.
IT Disaster recovery:
- Participating in disaster Recovery Governance & Planning
- Maintain and periodically review the inventory of critical business and IT services used for Disaster Recovery planning and testing.
- Manage, annual Disaster Recovery and High Availability testing calendar.
- Ensure DR plans align with business continuity requirements and regulatory obligations.
- Ensure ITSCM plans, risks, controls, and activities remain aligned with business continuity requirements.
- Assess continuity risks and recommend mitigation strategies.
- Lead end-to-end Disaster Recovery and High Availability testing activities.
- Coordinate pre-test planning, stakeholder communications, execution workshops, and test rehearsals.
- Manage test windows, bridge calls.
- Monitor runbook execution and ensure successful completion of recovery activities.
- Coordinate evidence collection, issue tracking, risk logging, and post-test reviews.
- Prepare executive-level DR testing reports and recommendations
- Facilitate coordination between technical teams, business stakeholders, management, and external vendors.
- Participate in crisis management and business continuity response activities.
Availability Management:
- Execute and govern the Availability Management process in line with agreed SLAs, OLAs, and service commitments.
- Establish and maintain service availability baselines and performance targets.
- Conduct regular service availability reviews with stakeholders.
- Monitor availability, uptime, and downtime of all applications and infrastructure services utilizing approved monitoring tools.
- Validate availability calculations and ensure data integrity across reporting platforms.
- Generate daily, weekly, and executive-level availability reports.
- Analyze service outages, performance degradation, availability trends, and recurring failures.
- Identify risks affecting service availability and develop mitigation plans.
- Recommend service reliability and resilience improvements.
- Validate exclusion of approved downtime from SLA calculations.
- Conduct service review meetings and availability governance discussions.
- Define monitoring requirements, availability targets, reporting standards, and governance controls for new services.
- Identify opportunities to improve service availability and operational efficiency.
- Drive root cause trending and proactive improvement initiatives.
- Submit recommendations to CSI governance boards.
- Measure effectiveness of implemented improvements.
- Enhance availability analytics using Power BI, ServiceNow, Dynatrace, SiteScope, SCOM, SolarWinds, or equivalent monitoring tools.
- Reduce manual effort through automation of KPI reporting and trend analysis.
Required Skill:
- Disaster Recovery Management
- Business Continuity Management (BCM)
- Business Impact Analysis (BIA)
- Availability Management
- RTO/RPO Management
- DR Runbook Management
- Power BI, ServiceNow (User level)
- Dynatrace, SiteScope, SolarWinds, (User level)
- Executive Reporting & Stakeholder Management.
KPI:
- DR and HA testing completion rate.
- DR testing success rate.
- Achievement of agreed RTO and RPO targets.
- Closure rate of DR findings and lessons learned.
- DR readiness compliance score.
- IT service availability percentage.
- SLA compliance achievement.
- Completion of availability improvement initiatives.
- Automation and reporting efficiency improvements
رقم الإعلان لدى المصدر: 4476576793