جامعة الملك عبدالله للعلوم والتقنية تعلن عن وظيفة قائد أنظمة الحوسبة عالية الأداء في السعودية
تفاصيل الوظيفة
تعلن جامعة الملك عبدالله للعلوم والتقنية (KAUST) عن توفر وظيفة مسؤول أنظمة الحوسبة عالية الأداء الرئيسي (HPC Lead Systems Administrator) في المملكة العربية السعودية.
نبذة عن الوظيفة
يتولى شاغل الوظيفة قيادة فريق يضمن التشغيل السلس للكتلة (Cluster) من أنظمة لينكس التي تضم أكثر من 300 عقدة حوسبة (GPU/CPU) إلى جانب أنظمة الملفات المتوازية والشبكات عالية الأداء. يجمع الدور بين الجوانب التقنية والقيادية، مع الإشراف على 3-4 من مسؤولي أنظمة HPC ذوي الخبرة، وتطوير وتنفيذ إجراءات التشغيل القياسية للنظام والفريق.
المهام والمسؤوليات
- تخطيط تشغيل النظام وترقيته لتلبية متطلبات المختبر والعملاء
- تطوير وتنفيذ سياسات جدولة أعباء العمل (Workload Scheduler)
- دعم أنظمة الملفات عالية الأداء
- إدارة البنية التحتية للشبكات بما يشمل TCP/IP وشبكات HPC
- استخدام لغات البرمجة النصية لأتمتة العقد وإدارة التهيئة
- إدارة أعطال الأجهزة وقطع الغيار
- بناء علاقات فعالة مع الموظفين وأعضاء هيئة التدريس والطلاب عبر Core Labs
- إدارة مشاريع متعددة أو مهمة تستخدم تقنيات تخطيط مشاريع متقدمة
- تخطيط وجدولة وتنفيذ أو تنسيق مراحل تفصيلية من عمل مشروع رئيسي أو مشروع كامل متوسط النطاق
- تحديد احتياجات التدريب التقني للموظفين في المجال
- المشاركة كعضو في الاستجابة لحوادث الأمن والسلامة
- خلق فرص لتحسين المنهجية التقنية أو المحتوى من خلال توسيع الجهود الحالية أو تطوير جهود جديدة؛ قد يوسع التكنولوجيا في مجالات تطبيقية جديدة؛ يسهم أو يقود أنشطة التطوير الفكري الرئيسية
- تقديم حلول مبتكرة لحل المشكلات لتعزيز قدرات المؤسسة؛ استخدام شبكة الأقران لتوسيع القدرات التقنية وتحديد فرص بحثية جديدة
- فهم الأهداف الاستراتيجية العامة والمساهمة فيها؛ رعاية والحفاظ على العلاقات مع كبار العملاء
- بدء مفاهيم مشاريع جديدة؛ تطوير مقترحات تقنية وتقديم عروض للعملاء المحتملين
- الإشراف على العديد من العلماء أو المهندسين أو الفنيين في العمل المكلف به؛ تقديم مساهمة رئيسية في تزويد فرق المشروع بالكوادر؛ بناء فرق وموظفين لتحسين الكفاءة والفعالية من حيث التكلفة
- تحديد وتقييم المرشحين للوظائف الشاغرة؛ توجيه/تدريب الموظفين في تطوير المهارات التقنية ومهارات المشاريع وتطوير الأعمال
الشروط والمتطلبات
درجة البكالوريوس في العلوم (أو ما يعادلها) في تخصص ذي صلة مع 10 سنوات خبرة، أو درجة الماجستير في العلوم (أو ما يعادلها) في تخصص ذي صلة مع 7 سنوات خبرة، أو درجة الدكتوراه (أو ما يعادلها) في تخصص ذي صلة مع 5 سنوات خبرة.
المهارات المطلوبة
- مدير عبء العمل SLURM بما في ذلك جدولة GPU
- أنظمة الملفات المتوازية (Weka IO, Lustre)
- شبكات TCP/IP والشبكات عالية الأداء (Infiniband)
- إتقان لغات البرمجة النصية (مثل Bash وPython وRuby)
- الإلمام بأدوات إدارة التهيئة (Puppet)
- مهارات توثيق متقنة
- القدرة على التعامل مع المستخدمين والموردين
- اتباع نهج تحليلي ومنهجي في حل المشكلات
- المبادرة في تحديد وترتيب فرص التطوير المناسبة
- مهارات اتصال فعالة باللغة الإنجليزية كتابةً وتحدثاً
- العمل بفعالية مع الفرق الأخرى في مختبر الحوسبة الفائقة
- تخطيط وجدولة ومراقبة العمل الخاص (وعمل الآخرين) بكفاءة ضمن مواعيد نهائية محددة ووفقاً للتشريعات والإجراءات ذات الصلة
- القدرة على العمل بنجاح في بيئة بحثية تعاونية عالية
- استخدام التقدير في تحديد وحل المشكلات والمهام المعقدة
- أداء مجموعة واسعة من الأعمال، أحياناً معقدة وغير روتينية، في بيئات متنوعة
- الحفاظ على معرفة على مستوى الخبراء في معظم أنظمة المختبر، بما يشمل إدارة أنظمة الحوسبة عالية الأداء، أو إدارة التخزين عالي الأداء، أو إدارة الشبكات عالية الأداء
عرض النص الأصلي للإعلان
Position Summary
Serve as the Lead for the team ensuring smooth operation of the Linux cluster consisting of 300+ GPU/CPU compute nodes including parallel filesystems and high-performance network. This is partly technical and partly people leading role which involves supervision of 3-4 experienced HPC system administrators. The role involves development, implementation and supervision of standard operating procedures for the system and the team.
Major Responsibilities
- System operation and upgrade planning to meet laboratory and customer requirements
- Workload scheduler policy development and implementation
- Support of high-performance filesystems
- Network infrastructure management including TCP/IP and HPC networks
- Use of scripting languages for nodes automation and configuration management
- Hardware failures and spare part management
- Build effective relationships with staff, faculty and students through the Core Labs
- Manages multiple or significant projects which may require the use of sophisticated project planning techniques
- Plans, schedules, conducts, or coordinates detailed phases of the work of a major project or in a total project of moderate scope
- Identifies technical training needs for staff attached to the area
- Serve as a resource and as a member to respond to security and safety incidents
- Creates opportunities to enhance technical methodology or content through expansion of existing, or development of, new efforts; may extend technology into new application areas; contributes or leads in major intellectual development activities
- Provides innovative problem-solving approaches to enhance organizational capabilities; uses peer network to expand technical capabilities and identify new research opportunities
- Understands broad strategic objectives and contributes to them; nurtures and maintains relationships with major customers
- May initiate new project concepts; develops technical proposals and makes presentations to potential customers
- Will supervise several scientists, engineers or technicians on assigned work; provides major input to staffing of overall project teams; builds teams and staff to optimize efficiency and cost effectiveness
- Identifies and evaluates candidates for open positions; mentors/trains staff in development of technical, project and business development skills
Competencies
- SLURM workload manager including GPU scheduling
- Parallel filesystems (Weka IO, Lustre)
- TCP/IP and high performance networks (Infiniband)
- Proficient in scripting languages (i.e. Bash, Python, Ruby)
- Familiar with configuration management tools (Puppet)
- Proficient documentation skills
- Will have working level contact with users and suppliers
- Demonstrates an analytical and systematic approach to problem solving
- Takes the initiative in identifying and negotiating appropriate development opportunities
- Demonstrates effective communication skills in written and oral English
- Works effectively with other teams in the Supercomputing Laboratory
- Plans, schedules and monitors own work (and that of others) competently within limited deadlines and according to relevant legislation and procedures
- Ability to work successfully in a highly collaborative research environment
- Uses discretion in identifying and resolving complex problems and assignments
- Performs a broad range of work, sometimes complex and non-routine, in a variety of environments
- Maintain expert-level knowledge in most of the laboratory systems, including high performance computing systems administration, high performance storage administration, or high performance network administration
Qualifications and Experience
- Bachelor of Science (or equivalent) in a relevant discipline plus 10 years’ experience, OR Master of Science (or equivalent) in a relevant discipline plus 7 years’ experience OR Doctor of Philosophy (or equivalent) in a relevant discipline plus 5 years’ experience.
وظائف أخرى لدى جامعة الملك عبدالله للعلوم والتقنية
جامعة الملك عبدالله للعلوم والتقنية تعلن عن وظيفة قائد حوسبة مستخدمين بحثية في السعودية
جامعة الملك عبدالله للعلوم والتقنية تعلن عن وظيفة أخصائي حوسبة مستخدمين بحثية (Linux) في السعودية
جامعة الملك عبدالله للعلوم والتقنية تعلن عن وظيفة أخصائي سلامة الإشعاع في السعودية
جامعة الملك عبدالله للعلوم والتقنية تعلن عن وظيفة خبير مهندس/مطور تكامل في السعودية