سلة تعلن عن وظيفة مهندس أول لموثوقية الموقع SRE في مكة المكرمة
Senior Site Reliability Engineer (SRE)
🏢 سلة (Salla)
تفاصيل الوظيفة
تعلن شركة سلة عن توفر وظيفة مهندس موثوقية أول (Senior Site Reliability Engineer - SRE) في مكة المكرمة، المملكة العربية السعودية. يفضل المرشحون المقيمون في المناطق الزمنية GMT 0 إلى +6 لتوافق فريق العمل وتغطية النوبات.
المهام والمسؤوليات
- قيادة الاستجابة للحوادث عالية الخطورة وإجراء مراجعات ما بعد الحادث.
- استكشاف المشكلات المعقدة وإصلاحها عبر التطبيقات والبنية التحتية والشبكات.
- تحسين متوسط وقت الإصلاح (MTTR) من خلال تحسين المراقبة والتنبيهات وأدوات التشخيص.
- المشاركة في نظام النوبات (on-call) لدعم أنظمة الإنتاج.
- تحديد وحل اختناقات الأداء وتحديات التوسع.
- إجراء اختبارات الحمل وتخطيط السعة لسيناريوهات الحركة العالية.
- تحسين البنية التحتية السحابية وعمليات النشر والأتمتة.
- تحسين آليات المرونة وتحمل الأخطاء والتعافي عبر الأنظمة.
- بناء وتحسين لوحات المعلومات والتنبيهات والمقاييس والسجلات والتتبعات.
- تحديد مؤشرات SLIs وأهداف SLOs وتحسين رؤية سلوك النظام.
- تطوير أدوات تقلل من الجهد التشغيلي وتزيد الموثوقية.
- المساهمة في البنية التحتية ككود (IaC) وخطوط CI/CD وسير عمل GitOps.
- العمل عن كثب مع فرق الهندسة لضمان أن الخدمات قوية وجاهزة للإنتاج.
- توجيه المهندسين في ممارسات الموثوقية وتصحيح الأخطاء والتشغيل.
الشروط والمتطلبات
- خبرة قوية في Kubernetes وتقنيات service mesh ومنصات السحابة (AWS أو GCP أو Azure).
- فهم عميق لأنظمة Linux والشبكات والأنظمة الموزعة وتوزيع الأحمال.
- خبرة عملية في Terraform أو أدوات البنية التحتية ككود (IaC) المشابهة.
- خبرة في منصات المراقبة مثل Prometheus أو Grafana أو Loki أو Mimir أو Elastic أو ما يعادلها.
- إتقان لغات البرمجة النصية أو البرمجة مثل Bash أو Python أو Go.
- خبرة في خطوط CI/CD وممارسات GitOps.
- مهارات قوية في تصحيح الأخطاء والاستجابة للحوادث وتحليل الأداء.
المهارات المطلوبة
- خلفية في الأنظمة كبيرة الحجم وعالية الحركة.
- خبرة في التصميم القابل لتحمل الأخطاء واستعادة الكوارث (DR) وأنماط التوفر العالي (HA).
- الإلمام بمفاهيم SLOs وSLIs وميزان الأخطاء (error budgets).
عرض النص الأصلي للإعلان
As a Senior SRE at Salla, you will lead reliability initiatives, handle complex incidents, improve platform performance, and guide engineering teams toward building resilient systems. You will also participate in the on-call rotation as part of our commitment to platform reliability.
Reliability & Incident Management
- Lead high-severity incident response and drive post-incident reviews.
- Troubleshoot complex issues across applications, infrastructure, and networks.
- Improve MTTR through better monitoring, alerts, and diagnostic tooling.
- Participate in the on-call rotation supporting production systems.
Performance & Scalability
- Identify and resolve performance bottlenecks and scaling challenges.
- Conduct load testing and capacity planning for high-traffic scenarios.
Infrastructure & Operations
- Enhance cloud-native infrastructure, deployment processes, and automation.
- Improve resilience, fault-tolerance, and recovery mechanisms across systems.
Observability
- Build and refine dashboards, alerts, metrics, logs, and traces.
- Define SLIs/SLOs and improve visibility into system behavior.
Tooling & Automation
- Develop tools that reduce operational toil and increase reliability.
- Contribute to infrastructure-as-code, CI/CD pipelines, and GitOps workflows.
Collaboration
- Work closely with engineering teams to ensure services are robust and production-ready.
- Mentor engineers on reliability, debugging, and operational best practices.
Bonus Skills
- Background in large-scale, high-traffic systems.
- Experience with fault-tolerant design, DR, and HA patterns.
- Familiarity with SLOs, SLIs, and error budgets.
Location Preference
- Candidates located within GMT 0 to +6 time zones are preferred to align with team collaboration and on-call coverage.
Requirements
- Strong experience with Kubernetes, service mesh technologies, and cloud platforms (AWS, GCP, or Azure).
- Deep understanding of Linux, networking, distributed systems, and load balancing.
- Hands-on experience with Terraform or similar Infrastructure-as-Code tools.
- Experience with observability platforms such as Prometheus, Grafana, Loki, Mimir, Elastic, or equivalent.
- Proficiency in scripting or programming languages such as Bash, Python, or Go.
- Experience with CI/CD pipelines and GitOps practices.
- Strong debugging, incident response, and performance analysis skills.
المصدر: الموقع الرسمي للجهة - أُضيفت للموقع في 10 يونيو 2026
وظائف أخرى لدى سلة