📍 المملكة العربية السعودية تحديث مستمر على مدار الساعة

سلة تعلن عن وظيفة عمليات البيانات والتعلم الآلي في المدينة المنورة

Data & ML Ops
🏢 سلة (Salla)
🕒 نُشرت: (منذ 9 أشهر) 📍 المدينة المنورة وظائف الهندسة والتقنية
التقديم على الوظيفة من المصدر الرسمي ↗

تفاصيل الوظيفة

تسعى شركة سلة إلى تعيين مهندس موثوقية مواقع أول (Senior Site Reliability Engineer) للعمل في المدينة المنورة، للمساهمة في تصميم وتوسيع نطاق وتأمين البنية التحتية للمنصة سريعة النمو، مع ضمان التوافر والأداء وكفاءة التكلفة على نطاق واسع.

المهام والمسؤوليات

  • تصميم ونشر ومراقبة وصيانة أحمال العمل الإنتاجية عبر مجموعات Kubernetes (EKS/AKS/GKE).
  • بناء أنظمة ذاتية الشفاء والتوسع التلقائي لتقليل التدخل اليدوي وضمان وقت التشغيل المستمر.
  • تحسين الشبكات والتحكم في حركة الدخول/الخروج و Service Mesh لضمان اتصال آمن وفعال.
  • تصميم وتشغيل منصات قواعد بيانات وتخزين موثوقة (SQL و NoSQL ومخازن الكائنات) داخل بيئات Kubernetes.
  • إدارة استراتيجيات النسخ الاحتياطي والتعافي من الكوارث والتكرار والتبديل الاحتياطي لتحقيق أهداف RPO/RTO.
  • تحسين أداء التخزين وتكلفته من خلال استراتيجيات متعددة المستويات وفصل البيانات الساخنة/الباردة وسياسات دورة حياة S3/التفريغ.
  • استكشاف مشكلات Kubernetes Persistent Volumes وحلها أثناء الحوادث (StorageClasses و CSI drivers و PVC issues).
  • تأمين وتوسيع نطاق منصات تخزين الكائنات (مثل MinIO/S3) ودمجها مع أحمال العمل لخطوط أنابيب بيانات عالية الإنتاجية.
  • العمل مع التخزين القائم على الكتل (EBS/io2/gp3) وأنظمة الملفات المشتركة (EFS و NFS) لتحقيق التوازن بين الأداء والمرونة والتكلفة.
  • دعم أفضل ممارسات GitOps و CI/CD باستخدام أدوات مثل ArgoCD و Flux و GitHub Actions.
  • بناء أتمتة لتوفير البنية التحتية وترقيتها باستخدام Terraform و Helm و Kubernetes Operators.
  • تقليل مخاطر الإصدار من خلال استراتيجيات التوزيع التدريجي (blue/green و canary و spot instance rolling updates).
  • إدارة مجموعة المراقبة والتنبيه (Prometheus و Grafana و Loki و VictoriaMetrics و OpenSearch).
  • قيادة إدارة الحوادث ومراجعات ما بعد الحادث لمنع تكرارها.
  • توفير رؤية فورية لحالة النظام وأدائه ومقاييس التكلفة.
  • تنفيذ سياسات IAM بأقل الامتيازات وتأمين الاتصال بين الخدمات وقوائم التحكم في الوصول (ACLs) وجدران الحماية.
  • فرض Kubernetes RBAC وإدارة الأسرار وسلسلة توريد الصور الآمنة.
  • المشاركة في الاستعداد للتدقيق وجهود الامتثال.
  • تحليل وضبط أداء النظام تحت أحمال عالية (CPU/ذاكرة/إدخال/إخراج).
  • التعاون مع فرق المنتج والمنصة لتحسين حجم الكتل وقواعد البيانات وطبقات التخزين.
  • تقديم لوحات معلومات لرؤية التكلفة لقيادة الهندسة.

الشروط والمتطلبات

  • 8+ سنوات من الخبرة في أدوار SRE / DevOps / Infrastructure Engineering.
  • خبرة عميقة في Kubernetes (متعدد المجموعات، تطوير Helm chart، شبكات متقدمة).
  • إتقان سير عمل GitOps باستخدام ArgoCD/Flux.
  • خبرة في AWS (يفضل) أو Azure/GCP، بالإضافة إلى البنية التحتية كرمز (Terraform, Pulumi, CloudFormation).
  • معرفة متقدمة بقواعد البيانات SQL و NoSQL (MySQL/Aurora, PostgreSQL, MongoDB, Redis).
  • مهارات البرمجة/الأتمتة في Python أو Bash أو Go.
  • خلفية قوية في المراقبة/الرصد (Prometheus, Grafana, Loki, ELK/Opensearch, VictoriaMetrics).
  • خبرة في CI/CD على نطاق واسع وإدارة حوادث الإنتاج.
  • خبرة في البث/المراسلة (Kafka, RabbitMQ أو ما يشابهها).

المهارات المطلوبة

  • خبرة في إدارة الأنظمة الحرجة على نطاق واسع (حركة مرور عالية، متعدد المناطق).
  • إثبات تحسين التكلفة في بيئات السحابة/Kubernetes.
  • الإلمام بـ Service Mesh (Istio, Linkerd) أو الشبكات المتقدمة/التحكم في الخروج.
  • خبرة مع مكونات منصة البيانات (Airflow, Debezium, ClickHouse, إلخ) تعتبر ميزة إضافية.
  • مهارات تواصل قوية والعمل الجماعي - القدرة على التعاون عبر فرق الهندسة و DevOps والأمان والمنتج.

المزايا

  • برامج تدريب وتطوير شاملة.
  • حوافز مكافآت قائمة على الأداء.
  • خيارات عمل مرنة من المنزل.
عرض النص الأصلي للإعلان

We are looking for a Senior Site Reliability Engineer (SRE) to help design, scale, and secure our rapidly growing platform infrastructure.
You will work across all critical systems - from customer-facing applications and APIs to internal platforms and data services - ensuring availability, performance, and cost efficiency at scale.

You’ll be hands-on with Kubernetes, observability, GitOps, automation, and cloud infrastructure, while partnering closely with application, platform, and data teams to deliver a highly reliable and self-healing environment.

This role is ideal for an engineer who thrives on complex distributed systems, loves to automate everything, and can balance speed, stability, and cost-efficiency in production.

  • Bachelor’s degree in Computer Science, Engineering, or a related field - or equivalent work experience.
  • Design, deploy, monitor, and maintain production workloads across Kubernetes (EKS/AKS/GKE) clusters.
  • Build self-healing, auto-scaling systems that minimize manual intervention and ensure uptime.
  • Design and operate reliable database and storage platforms (SQL, NoSQL, and object stores) within Kubernetes environments.
  • Implement backup, disaster recovery, replication, and failover strategies to meet RPO/RTO targets.
  • Troubleshoot and recover Kubernetes Persistent Volumes (StorageClasses, CSI drivers, PVC issues).
  • Optimize storage performance and cost through multi-tier strategies, hot/cold data separation, and S3/offloading lifecycle policies.
  • Secure and scale object storage platforms (e.g., MinIO/S3-compatible) for high-throughput data pipelines.
  • Manage block storage (EBS/io2/gp3) and shared file systems (EFS, NFS) for resilience and cost balance.
  • Collaborate with teams to optimize networking, ingress/egress traffic, and service mesh for secure communication.

Platform & Infrastructure Reliability

  • Design, deploy, monitor, and maintain production workloads across Kubernetes (EKS/AKS/GKE) clusters.
  • Build self-healing, auto-scaling systems that minimize toil and manual intervention.
  • Optimize networking, ingress/egress traffic control, and service mesh for secure & performant communication.
  • Design and operate reliable database and storage platforms (SQL, NoSQL, and object stores) in Kubernetes environments.
  • Own backup, disaster recovery, replication, and failover strategies to meet RPO/RTO targets for critical data services.
  • Optimize storage performance and cost through multi-tier strategies, hot/cold data separation, and S3/offloading lifecycle policies.
  • Troubleshoot and recover Kubernetes Persistent Volumes confidently during incidents (StorageClasses, CSI drivers, PVC issues).
  • Secure and scale object storage platforms (e.g., MinIO/S3-compatible) and integrate with workloads for high-throughput data pipelines.
  • Work with block storage (EBS/io2/gp3) and shared file systems (EFS, NFS) to balance performance, resiliency, and cost.

Automation & Delivery

  • Champion GitOps and CI/CD best practices (ArgoCD, Flux, GitHub Actions).
    Build automation for infrastructure provisioning and upgrades using Terraform, Helm, and Kubernetes Operators.
  • Reduce release risk through progressive delivery strategies (blue/green, canary, spot instance rolling updates).

Observability & Incident Response

  • Own the monitoring and alerting stack (Prometheus, Grafana, Loki, VictoriaMetrics, OpenSearch).
  • Lead incident management and postmortems to prevent recurrence.
  • Provide real-time visibility into system health, performance, and cost metrics.

Security & Compliance

  • Implement least-privilege IAM policies, secure service-to-service communication, and network ACLs/firewalls.
  • Enforce Kubernetes RBAC, secret management, and secure image supply chain.
  • Participate in audit readiness and compliance efforts.

Performance & Cost Optimization

  • Analyze and tune system performance under scale (CPU/memory/IO).
  • Partner with product and platform teams to right-size clusters, databases, and storage tiers.

Introduce cost visibility dashboards for engineering leadership.

Preferred Qualifications

  • Experience managing mission-critical systems at scale (high traffic, multi-region).
  • Proven cost optimization in cloud/K8s environments.
  • Familiarity with service mesh (Istio, Linkerd) or advanced networking/egress control.
  • Experience with data platform components (Airflow, Debezium, ClickHouse, etc.) is a plus but not required.

Strong communication skills and teamworker - able to collaborate across engineering, DevOps, security, and product teams.

Requirements

  • 8+ years in SRE / DevOps / Infrastructure Engineering roles.
  • Deep Kubernetes expertise (multi-cluster, Helm chart development, advanced networking).
  • Strong GitOps workflows using ArgoCD/Flux.
  • Expertise with AWS (preferred) or Azure/GCP, plus Infrastructure-as-Code (Terraform, Pulumi, CloudFormation).
  • Advanced knowledge of SQL & NoSQL databases (MySQL/Aurora, PostgreSQL, MongoDB, Redis).
  • Scripting/automation skills in Python, Bash, or Go.
  • Solid background in monitoring/observability (Prometheus, Grafana, Loki, ELK/Opensearch, VictoriaMetrics).
  • Experience with CI/CD at scale and managing production incidents.
  • Experience with streaming/messaging (Kafka, RabbitMQ, or similar).

Benefits

  • Comprehensive Training & Development programs.
  • Performance-based Bonus incentives.
  • Flexible Work From Home options.
المصدر: الموقع الرسمي للجهة - أُضيفت للموقع في 10 يونيو 2026

وظائف أخرى لدى سلة