📍 المملكة العربية السعودية تحديث مستمر على مدار الساعة

سلة تعلن عن وظيفة مهندس أول لموثوقية الموقع SRE في مكة المكرمة

Senior Site Reliability Engineer (SRE)
🏢 سلة (Salla)
🕒 نُشرت: (منذ 6 أشهر) 📍 مكة المكرمة وظائف الهندسة والتقنية
التقديم على الوظيفة من المصدر الرسمي ↗

تفاصيل الوظيفة

تعلن شركة سلة عن توفر وظيفة مهندس موثوقية أول (Senior Site Reliability Engineer - SRE) في مكة المكرمة، المملكة العربية السعودية. يفضل المرشحون المقيمون في المناطق الزمنية GMT 0 إلى +6 لتوافق فريق العمل وتغطية النوبات.

المهام والمسؤوليات

  • قيادة الاستجابة للحوادث عالية الخطورة وإجراء مراجعات ما بعد الحادث.
  • استكشاف المشكلات المعقدة وإصلاحها عبر التطبيقات والبنية التحتية والشبكات.
  • تحسين متوسط وقت الإصلاح (MTTR) من خلال تحسين المراقبة والتنبيهات وأدوات التشخيص.
  • المشاركة في نظام النوبات (on-call) لدعم أنظمة الإنتاج.
  • تحديد وحل اختناقات الأداء وتحديات التوسع.
  • إجراء اختبارات الحمل وتخطيط السعة لسيناريوهات الحركة العالية.
  • تحسين البنية التحتية السحابية وعمليات النشر والأتمتة.
  • تحسين آليات المرونة وتحمل الأخطاء والتعافي عبر الأنظمة.
  • بناء وتحسين لوحات المعلومات والتنبيهات والمقاييس والسجلات والتتبعات.
  • تحديد مؤشرات SLIs وأهداف SLOs وتحسين رؤية سلوك النظام.
  • تطوير أدوات تقلل من الجهد التشغيلي وتزيد الموثوقية.
  • المساهمة في البنية التحتية ككود (IaC) وخطوط CI/CD وسير عمل GitOps.
  • العمل عن كثب مع فرق الهندسة لضمان أن الخدمات قوية وجاهزة للإنتاج.
  • توجيه المهندسين في ممارسات الموثوقية وتصحيح الأخطاء والتشغيل.

الشروط والمتطلبات

  • خبرة قوية في Kubernetes وتقنيات service mesh ومنصات السحابة (AWS أو GCP أو Azure).
  • فهم عميق لأنظمة Linux والشبكات والأنظمة الموزعة وتوزيع الأحمال.
  • خبرة عملية في Terraform أو أدوات البنية التحتية ككود (IaC) المشابهة.
  • خبرة في منصات المراقبة مثل Prometheus أو Grafana أو Loki أو Mimir أو Elastic أو ما يعادلها.
  • إتقان لغات البرمجة النصية أو البرمجة مثل Bash أو Python أو Go.
  • خبرة في خطوط CI/CD وممارسات GitOps.
  • مهارات قوية في تصحيح الأخطاء والاستجابة للحوادث وتحليل الأداء.

المهارات المطلوبة

  • خلفية في الأنظمة كبيرة الحجم وعالية الحركة.
  • خبرة في التصميم القابل لتحمل الأخطاء واستعادة الكوارث (DR) وأنماط التوفر العالي (HA).
  • الإلمام بمفاهيم SLOs وSLIs وميزان الأخطاء (error budgets).
عرض النص الأصلي للإعلان

As a Senior SRE at Salla, you will lead reliability initiatives, handle complex incidents, improve platform performance, and guide engineering teams toward building resilient systems. You will also participate in the on-call rotation as part of our commitment to platform reliability.

Reliability & Incident Management

  • Lead high-severity incident response and drive post-incident reviews.
  • Troubleshoot complex issues across applications, infrastructure, and networks.
  • Improve MTTR through better monitoring, alerts, and diagnostic tooling.
  • Participate in the on-call rotation supporting production systems.

Performance & Scalability

  • Identify and resolve performance bottlenecks and scaling challenges.
  • Conduct load testing and capacity planning for high-traffic scenarios.

Infrastructure & Operations

  • Enhance cloud-native infrastructure, deployment processes, and automation.
  • Improve resilience, fault-tolerance, and recovery mechanisms across systems.

Observability

  • Build and refine dashboards, alerts, metrics, logs, and traces.
  • Define SLIs/SLOs and improve visibility into system behavior.

Tooling & Automation

  • Develop tools that reduce operational toil and increase reliability.
  • Contribute to infrastructure-as-code, CI/CD pipelines, and GitOps workflows.

Collaboration

  • Work closely with engineering teams to ensure services are robust and production-ready.
  • Mentor engineers on reliability, debugging, and operational best practices.

Bonus Skills

  • Background in large-scale, high-traffic systems.
  • Experience with fault-tolerant design, DR, and HA patterns.
  • Familiarity with SLOs, SLIs, and error budgets.

Location Preference

  • Candidates located within GMT 0 to +6 time zones are preferred to align with team collaboration and on-call coverage.

Requirements

  • Strong experience with Kubernetes, service mesh technologies, and cloud platforms (AWS, GCP, or Azure).
  • Deep understanding of Linux, networking, distributed systems, and load balancing.
  • Hands-on experience with Terraform or similar Infrastructure-as-Code tools.
  • Experience with observability platforms such as Prometheus, Grafana, Loki, Mimir, Elastic, or equivalent.
  • Proficiency in scripting or programming languages such as Bash, Python, or Go.
  • Experience with CI/CD pipelines and GitOps practices.
  • Strong debugging, incident response, and performance analysis skills.
المصدر: الموقع الرسمي للجهة - أُضيفت للموقع في 10 يونيو 2026

وظائف أخرى لدى سلة