سلة تعلن عن وظيفة محلل بيانات أول - AI (Gen AI & Recommendation Systems) في مكة المكرمة
تفاصيل الوظيفة
تسعى شركة سلة إلى تعيين محلل بيانات أول - ذكاء اصطناعي توليدي وأنظمة توصيات (Senior Data Analyst - AI (Gen AI & Recommendation Systems)) للعمل في مكة المكرمة، السعودية. يتطلب الدور خبرة في هندسة التحليلات وبناء خطوط البيانات، إلى جانب فهم قوي لسير عمل التعلم الآلي لدعم أنظمة التوصيات والذكاء الاصطناعي التوليدي.
المهام والمسؤوليات
- بناء وصيانة خطوط بيانات دفعية وموجهة بالتيار (batch & streaming pipelines) قابلة للتطوير ومتسامحة مع الأخطاء تدعم حالات الاستخدام التحليلي والتعلم الآلي.
- تحديد المقاييس الرئيسية لكل منتج يتم إطلاقه وبناء تقارير مركزية قوية حولها لإظهار الاتجاهات والرؤى المهمة.
- تصميم نماذج بيانات متعددة الطبقات (من raw إلى feature-ready marts) مع الحفاظ على الاتساق والأداء عبر نماذج ML ولوحات البيانات وواجهات API.
- هندسة تدفقات البيانات التي تغذي وتحدّث متجر الميزات الخاص بنماذج التعلم الآلي (Feature Store) - بما في ذلك بيانات الرسم البياني إذا لزم الأمر - مع ضمان التوفر وزمن الوصول المنخفض.
- بناء خطوط البيانات وأطر المقاييس الخاصة باختبارات A/B، بما في ذلك بنية experiment schema، وتسجيل التعيينات، وحساب المقاييس الموثوق للحصول على نتائج صحيحة إحصائياً.
- امتلاك قاعدة ClickHouse كخبير مجال: تصميم المخططات، ضبط الأداء، ودعم استعلامات سريعة لتجميع التجارب وتقديم الميزات.
- تنفيذ Change Data Capture (CDC) وتدفقات قائمة على الأحداث (مثل Apache Kafka) للحفاظ على تحديث البيانات حيثما تحتاج التقارير والتوصيات.
- بناء وإدارة سير العمل باستخدام أدوات orchestration حديثة (مثل Mage AI، Airflow، Prefect) لضمان تسليم موثوق وإدارة التبعيات.
- تعريف وتفسير المقاييس الصحيحة للترتيب offline/online، وهندسة الميزات التي تحتاجها النماذج فعلاً.
- التعاون مع علماء البيانات ومهندسي ML وفريق المنتج والفرق الخلفية لتحويل متطلبات البيانات إلى خطوط إنتاجية وميزات ML قابلة للتنفيذ.
الشروط والمتطلبات
- خبرة 4+ سنوات كمهندس بيانات/تحليلات في بناء أنظمة بيانات للتحليلات والتعلم الآلي.
- إتقان لغة Python و SQL متقدم.
- مهارات قوية في BI والتصور البياني (مثل Looker، Tableau) مع حدس جيد حول المقاييس المهمة وكيفية عرضها.
- خبرة عملية في بناء خطوط البيانات باستخدام أدوات orchestration حديثة (Mage AI، Airflow، Prefect).
- خبرة إنتاجية عميقة مع ClickHouse (أو BigQuery، Snowflake، أو ما شابه).
- خبرة عملية في نمذجة متعددة الطبقات (raw، staging، marts) باستخدام أنماط Kimball أو Data Vault أو OBT.
- فهم قوي لأطر التجارب - التعيين، holdouts، خطوط المقاييس، تقليل التباين.
- فهم جيد لدورة حياة ML - كيف تستهلك النماذج البيانات، وكيف تعمل متاجر الميزات (مثل Feast، Hopsworks)، وكيفية هندسة الميزات على نطاق واسع، بالإضافة إلى معرفة كافية بمقاييس الترتيب لدعم أنظمة التوصيات.
المهارات المطلوبة
- خبرة في DBT للنمذجة والتحويل.
- بناء أو دمج منصات A/B (مثل Statsig، Optimizely، GrowthBook، أو المخصصة).
- خبرة في Apache Kafka وأدوات CDC (مثل Debezium، Maxwell).
- خبرة في قواعد بيانات الرسم البياني (مثل Dgraph، Neo4j، Amazon Neptune) وهيكلة البيانات لها.
- معرفة بلغة JavaScript أو Go.
عرض النص الأصلي للإعلان
We're looking for a Senior Data Analyst/ Analytics Engineer to own data and analytics across our Gen AI and Recommendation Systems work. It's a hybrid role: you'll own the centralized reporting that turns data into decisions and build the pipelines and data models that feed it - defining the right metrics for each product we ship rather than waiting on others to prepare your data. For Recommendation Systems, you'll bring enough ML understanding to engineer the right features and evaluation metrics, partnering closely with Data Scientists, ML Engineers, Product, and Backend teams.
Key Responsibilities
- Pipeline Architecture & Development: Build and maintain scalable, fault-tolerant batch and streaming pipelines that serve analytical and ML use cases.
- Centralized Reporting & Metrics: Define the key metrics for each product we ship and build rock-solid centralized reporting around them, surfacing the trends and insights that matter.
- Data Modeling: Design and own multi-layer data models (staging to feature-ready marts) that stay consistent and performant across ML models, dashboards, and APIs, handling schema changes cleanly.
- Feature Store & ML Data Flows: Engineer the data flows that populate and update our ML Feature Store (and graph data where relevant) with the availability and low latency recommendation models need.
- Experimentation & A/B Testing: Build the pipelines and metrics frameworks behind A/B testing - experiment schemas, assignment logging, and reliable metric computation for statistically sound results.
- ClickHouse Mastery: Own ClickHouse as the domain expert - schema design, performance tuning, and fast queries for experiment aggregation and feature serving.
- Streaming & CDC: Implement Change Data Capture (CDC) and event-driven flows (e.g. Apache Kafka) to keep data fresh where reporting and recommendations need it.
- Orchestration & Automation: Build and manage workflows with modern orchestration tools (e.g. Mage AI, Airflow, Prefect) for reliable delivery and dependency management.
- ML-Aware Support: Define and interpret the right offline and online ranking metrics, and engineer the features the models actually need.
- Cross-Functional Collaboration: Partner with Data Scientists, ML Engineers, Product, and Backend to turn data requirements into production pipelines and actionable ML features.
Requirements
- Experience: 4+ years as a Data/Analytics Engineer building data systems for analytics and ML.
- Programming: Expert Python and advanced SQL.
- BI & Visualization: Strong BI/visualization skills (e.g. Looker, Tableau) and good intuition for which metrics matter and how to present them.
- Pipelines & Orchestration: Hands-on building pipelines with modern orchestration (Mage AI, Airflow, Prefect) - you build your own data, not just consume it.
- Data Warehouse / ClickHouse: Deep production experience with ClickHouse (or BigQuery, Snowflake, or similar).
- Data Modeling: Hands-on multi-layer modeling (raw, staging, marts) using Kimball, Data Vault, or OBT patterns.
- Experimentation & A/B Testing: Solid grasp of experimentation frameworks - assignment, holdouts, metric pipelines, variance reduction.
- ML Exposure: Good grasp of the ML lifecycle - how models consume data, how Feature Stores work (e.g. Feast, Hopsworks), and how to engineer features at scale, plus enough ranking-metric knowledge to support Recommendation Systems.
Nice to have:
- DBT for modeling and transformation.
- Building or integrating A/B platforms (e.g. Statsig, Optimizely, GrowthBook, or custom).
- Apache Kafka and CDC tools (e.g. Debezium, Maxwell).
- Graph Databases (e.g. Dgraph, Neo4j, Amazon Neptune) and structuring data for them.
- JavaScript or Go.
وظائف أخرى لدى سلة