📍 المملكة العربية السعودية تحديث مستمر على مدار الساعة

شركة TestCrew تعلن عن وظيفة مهندس Site Reliability في الأحساء السعودية

Site Reliability Engineer
🏢 TestCrew | شركة متخصصة في اختبار البرمجيات في مستوى الجودة والأداء
🕒 نُشرت: (اليوم) 📍 الأحساء وظائف الهندسة والتقنية
التقديم على الوظيفة من المصدر الرسمي ↗

تفاصيل الوظيفة

تسعى شركة TestCrew المتخصصة في اختبار البرمجيات وجودة الأداء إلى توظيف مهندسي موثوقية المواقع (Site Reliability Engineers) للعمل في موقع العميل بالأحساء (المنطقة الشرقية). سيتولى المرشحون الناجحون مسؤولية ضمان توافر وموثوقية وأداء وأمن الأنظمة الحيوية من خلال المراقبة الاستباقية والأتمتة والاستجابة المنضبطة للحوادث.

المهام والمسؤوليات

  • مراقبة وإدارة البنية التحتية والتطبيقات والخدمات لضمان التوافر العالي والاستقرار والأداء الأمثل.
  • تشغيل وصيانة منصات المراقبة والملاحظة لكشف الحوادث وتحليل الاتجاهات والاستجابة الاستباقية لمشكلات النظام.
  • تصميم وتنفيذ استراتيجيات المراقبة ولوحات المعلومات والتنبيهات ومقاييس الأداء لتحسين الرؤية التشغيلية.
  • إجراء تحليل السبب الجذري (RCA) للحوادث وتنفيذ إجراءات وقائية لتقليل المشكلات المتكررة.
  • أتمتة المهام التشغيلية والنشر وأنشطة الصيانة باستخدام أدوات البرمجة النصية وأتمتة البنية التحتية.
  • إدارة خطوط أنابيب CI/CD ودعم عمليات إدارة الإصدار لتمكين عمليات نشر برمجيات موثوقة وفعالة.
  • تحسين أداء النظام مع ضمان الامتثال لاتفاقيات مستوى الخدمة (SLAs) والمعايير التشغيلية وأفضل الممارسات.
  • دعم مبادرات استمرارية الأعمال والتعافي من الكوارث والمرونة التشغيلية.
  • تنفيذ أفضل ممارسات الأمن، بما في ذلك تعزيز النظام وإدارة التصحيح والإجراءات التشغيلية الآمنة.
  • التعاون مع فرق التطوير والبنية التحتية والأمن وعمليات تكنولوجيا المعلومات لحل المشكلات الفنية المعقدة.
  • صيانة الوثائق التشغيلية ودفاتر التشغيل ومقالات المعرفة وتقارير الحوادث.
  • المشاركة في إدارة الحوادث الكبرى ومراجعات ما بعد الحوادث ومبادرات التحسين المستمر.
  • تقديم الدعم عند الطلب والمشاركة في المناوبات الدورية حسب الحاجة.

الشروط والمتطلبات

  • درجة البكالوريوس في علوم الحاسب أو تكنولوجيا المعلومات أو الهندسة أو مجال ذي صلة.
  • خبرة عملية تتراوح بين 2-5 سنوات في هندسة موثوقية المواقع (SRE) أو DevOps أو عمليات البنية التحتية أو دعم الإنتاج.
  • خبرة في إدارة منصات المراقبة والملاحظة مثل Grafana, Prometheus, Datadog, Instana, Zabbix أو حلول مكافئة.
  • مهارات قوية في تحليل الأعطال وتحليل السبب الجذري عبر البنية التحتية والتطبيقات والشبكات.
  • خبرة عملية في البرمجة النصية والأتمتة باستخدام Python, Bash, PowerShell.
  • خبرة في أدوات الأتمتة والبنية التحتية كرمز (IaC) مثل Ansible و Terraform.
  • معرفة عملية بأدوات CI/CD مثل Jenkins و GitLab CI و Azure DevOps.
  • فهم جيد لاستمرارية الأعمال والتعافي من الكوارث وإدارة المخاطر التشغيلية.
  • معرفة أفضل ممارسات الأمن السيبراني، بما في ذلك تعزيز النظام والتحكم في الوصول وإدارة الثغرات وإدارة التصحيح.
  • الإلمام بعمليات ITSM (إدارة الحوادث والمشكلات والتغيير) وأدوات مثل ServiceNow أو Jira Service Management (ميزة إضافية).
  • إجادة اللغتين العربية والإنجليزية قراءة وكتابة وتحدثاً.

المهارات المطلوبة

  • هندسة موثوقية المواقع (SRE)
  • مراقبة النظام والملاحظة
  • إدارة الحوادث وتحليل السبب الجذري
  • مراقبة الأداء وتحسينه
  • DevOps و CI/CD
  • البنية التحتية كرمز (Terraform, Ansible)
  • الأتمتة والبرمجة النصية (Python, Bash, PowerShell)
  • إدارة خوادم Linux و Windows
  • أساسيات الشبكات
  • استمرارية الأعمال والتعافي من الكوارث
  • عمليات الأمن السيبراني
  • عمليات ITSM
  • ServiceNow / Jira Service Management
  • التوثيق ودفاتر التشغيل
  • عقلية ملكية قوية مع نهج استباقي يركز على الموثوقية
  • مهارات تحليلية وحل مشكلات ممتازة
  • القدرة على البقاء هادئاً ومنهجياً أثناء الحوادث الحرجة
  • مهارات تواصل وتعاون قوية
  • القدرة على العمل بفعالية في بيئة مؤسسية سريعة الخطى
  • الاستعداد للعمل في موقع الأحساء والمشاركة في الدعم عند الطلب أو المناوبات الدورية
عرض النص الأصلي للإعلان
About the Role

TestCrew is seeking Site Reliability Engineers (SREs) to join an upcoming enterprise engagement in Al Ahsa. The successful candidates will be responsible for ensuring the availability, reliability, performance, and security of critical IT systems through proactive monitoring, automation, and disciplined incident response.

This is a hands-on operational role that requires close collaboration with development, operations, and security teams to maintain highly available services, streamline operations through automation, and drive continuous service improvement. Candidates must be willing to work onsite at the client location in Al Ahsa and participate in on-call or shift rotations as required.

Key Responsibilities
  • Monitor and manage enterprise infrastructure, applications, and services to ensure high availability, stability, and optimal performance.
  • Operate and maintain monitoring and observability platforms to detect incidents, analyze trends, and respond proactively to system issues.
  • Design and implement monitoring strategies, dashboards, alerts, and performance metrics to improve operational visibility.
  • Perform root cause analysis (RCA) for incidents and implement preventive measures to reduce recurring issues.
  • Automate operational tasks, deployments, and maintenance activities using scripting and infrastructure automation tools.
  • Manage CI/CD pipelines and support release management processes to enable reliable and efficient software deployments.
  • Optimize system performance while ensuring compliance with service level agreements (SLAs), operational standards, and best practices.
  • Support business continuity, disaster recovery, and operational resilience initiatives.
  • Implement security best practices, including system hardening, patch management, and secure operational procedures.
  • Collaborate with development, infrastructure, security, and IT operations teams to resolve complex technical issues.
  • Maintain operational documentation, runbooks, knowledge articles, and incident reports.
  • Participate in major incident management, post-incident reviews, and continuous improvement initiatives.
  • Provide on-call support and participate in shift rotations as required.
Required Qualifications
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
  • 2-5 years of hands-on experience in Site Reliability Engineering (SRE), DevOps, Infrastructure Operations, or Production Support.
  • Experience managing enterprise monitoring and observability platforms such as:
  • Grafana
  • Prometheus
  • Datadog
  • Instana
  • Zabbix
  • or equivalent solutions
  • Strong troubleshooting and root cause analysis skills across infrastructure, applications, and networking.
  • Hands-on experience with scripting and automation using:
  • Python
  • Bash
  • PowerShell
  • Experience with automation and Infrastructure as Code (IaC) tools such as:
  • Ansible
  • Terraform
  • Practical knowledge of CI/CD tools such as:
  • Jenkins
  • GitLab CI
  • Azure DevOps
  • Good understanding of business continuity, disaster recovery, and operational risk management.
  • Knowledge of cybersecurity best practices, including system hardening, access control, vulnerability management, and patch management.
  • Familiarity with ITSM processes (Incident, Problem, and Change Management) and tools such as ServiceNow or Jira Service Management is an advantage.
  • Fluent in both Arabic and English, with strong written and verbal communication skills.
Preferred Qualifications
  • ITIL Foundation Certification.
  • Cloud platform experience (AWS, Azure, or Google Cloud Platform).
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes.
  • Experience supporting large-scale enterprise or government environments.
Technical Skills
  • Site Reliability Engineering (SRE)
  • System Monitoring & Observability
  • Incident Management & Root Cause Analysis
  • Performance Monitoring & Optimization
  • DevOps & CI/CD
  • Infrastructure as Code (Terraform, Ansible)
  • Automation & Scripting (Python, Bash, PowerShell)
  • Linux & Windows Server Administration
  • Networking Fundamentals
  • Business Continuity & Disaster Recovery
  • Cybersecurity Operations
  • ITSM Processes
  • ServiceNow / Jira Service Management
  • Documentation & Operational Runbooks
Personal Attributes
  • Strong ownership mindset with a proactive, reliability-first approach.
  • Excellent analytical and problem-solving skills.
  • Ability to remain calm and methodical during critical incidents.
  • Strong communication and collaboration skills.
  • Ability to work effectively in a fast-paced enterprise environment.
  • Willingness to work onsite in Al Ahsa and participate in on-call or shift support as required.

  • المصدر: LinkedIn - أُضيفت للموقع في 4 أغسطس 2026

    وظائف أخرى لدى TestCrew | شركة متخصصة في اختبار البرمجيات في مستوى الجودة والأداء