ملاءة تعلن عن وظيفة عالم بيانات أول - ذكاء المعاملات في الرياض
تفاصيل الوظيفة
شركة ملاءة (Malaa)، أول منصة مصرفية مفتوحة للبيع بالتجزئة في المملكة العربية السعودية، تعلن عن وظيفة Senior Data Scientist - Transaction Intelligence في الرياض.
نبذة عن الوظيفة
تمثل هذه الوظيفة المحور الأساسي لتحويل مئات الملايين من سجلات البيانات السلوكية الخام إلى رؤى قابلة للتنفيذ. تمتد مسؤولياتها عبر كامل خط أنابيب البيانات: إثراء المعاملات (تصنيف سلاسل نصية قصيرة ومزعجة للمتاجر بخليط من الإنجليزية والعربية المنقولة بالحروف اللاتينية، وتحديد الكيانات التجارية الحقيقية)، ثم بناء النماذج والتحليلات فوقها لتحويل بيانات المعاملات إلى رؤى عبر خدمات الشركة - التصنيف، التعرف على الأنماط، كشف الحالات الشاذة، وتقييم المخاطر. ستعمل كعالم بيانات لبيانات المعاملات عبر الشركة بأكملها: امتلاك نماذج الإثراء من البداية إلى النهاية (النمذجة، التقييم، وكتابة كود الإنتاج)، والتعاون مع فرق المنتج والهندسة لتصميم وإطلاق وقياس الميزات المعتمدة على النماذج في الخدمات الأخرى.
المهام والمسؤوليات
- امتلاك ورفع مستوى إثراء المعاملات - الأساس الذي يقوم عليه كل شيء آخر: تصنيف يعمم على متاجر لم نرها من قبل، ومطابقة أسماء المتاجر التي تتحمل الاختلافات غير الضارة دون الخلط بين الأنشطة التجارية المختلفة فعلياً.
- امتلاك ثقة النموذج عبر أنظمتنا: احتمالات معايرة جيداً، وامتناع مبدئي عن الحالات غير المؤكدة، وتوجيه قائم على الثقة تعتمد عليه المنتجات وسير العمل المرجعي.
- بناء نماذج رؤى مبنية على المعاملات للخدمات الأخرى: التعرف على الأنماط، كشف الحالات الشاذة، والتحليل السلوكي على تيارات المعاملات.
- امتلاك التحقق المستمر من نماذج المخاطر لدينا: اختبار رجعي للنتائج مقابل السلوك الفعلي للمستخدمين، مراقبة التمييز والمعايرة مع تطور السلوكيات، ودفع تحسينات النموذج بناءً على ما تظهره البيانات.
- معالجة بيانات النصوص والسلوك كما هي فعلاً: ثنائية اللغة، عربية رومنة غير رسمية مع تهجئة غير مستقرة؛ تيارات معاملات مع آثار اقتطاع وخصائص بنكية وتوزيعات ثقيلة الذيل.
- بناء خطوط أنابيب دفعية فعالة على نطاق مئات الملايين من السجلات، مصممة لإعادة التشغيل دورياً مع تحسن النماذج.
- التصميم ضمن حدود شركة مالية منظمة: حوكمة البيانات، الخصوصية، ومتطلبات الأمن السيبراني تشكل البيانات التي يمكننا استخدامها وأين يمكن تشغيل أعباء العمل والأدوات التي يمكننا اعتمادها - ستبني حلولاً ممتازة داخل هذه الحدود، بالعمل مع تلك الفرق وليس حولهم.
- بناء الانضباط التقييمي الذي تستحقه هذه الأنظمة: مجموعات بيانات موسومة، مجموعات اختبار انحدار لسلوك النموذج، ومقاييس تجيب على سؤال 'هل ساعد هذا التغيير؟' لكل إصدار.
الشروط والمتطلبات
- 3+ سنوات من الخبرة التطبيقية في تعلم الآلة/علوم البيانات، مع نماذج تم شحنها وصيانتها في الإنتاج على نطاق واسع - ومعرفة بأنماط فشلها.
- نطاق تحليلي قوي يتجاوز النمذجة: تحليل استكشافي، دقة إحصائية، وتصميم ميزات على بيانات سلوكية/جدولية (إجادة SQL مفترضة).
- خبرة عملية في تشابه النصوص، المطابقة التقريبية، أو حل الكيانات على السلاسل النصية الواقعية المزعجة.
- إجادة قوية للغة Python والمكدس العلمي (scikit-learn, scipy, numpy)، مع وعي بالأداء والتكلفة على نطاق واسع - التوجيه، هياكل البيانات المتناثرة، الحساب الدفعي الفعال.
- نضج مثبت في العمل ضمن قيود مفروضة خارجياً - حوكمة البيانات، الخصوصية، الأمن السيبراني، الامتثال، سياسات البنية التحتية - بما في ذلك التعاون مع الفرق التي تملكها.
- سجل حافل في تحويل شكاوى غامضة مثل 'النموذج يبدو خاطئاً' إلى خصائص مقاسة ومختبرة بالانحدار.
- الراحة في قراءة وتصحيح كود النموذج الذي لم تكتبه، والعمل مباشرة مع فرق المنتج على مشكلات غير محددة بوضوح.
المهارات المطلوبة
- معالجة النصوص العربية / العربيزي - ميزة قوية؛ بياناتنا ثنائية اللغة مع رومنة غير مستقرة.
- نمذجة وتقييم مخاطر الائتمان أو السلوك - بطاقات الأداء، قياس التمييز والمعايرة، الاختبار الرجعي مقابل النتائج المحققة.
- تضمين النصوص واسترجاع الجيران التقريبيين في بيئات محدودة الموارد.
- خطوط أنابيب التوسيم بمساعدة LLM، التقطير، أو الإشراف الضعيف.
- أدوات دورة حياة تعلم الآلة وأدوات الخدمة الدفعية (تتبع التجارب، إصدار البيانات، قوائم المهام الموزعة).
عرض النص الأصلي للإعلان
About the role:
Malaa is Saudi Arabia’s first retail Open Banking platform in the Kingdom. Our product is built
on one core asset: hundreds of millions of records of behavioral data. Turning that raw data into
something people can act on is this role's whole job, and it spans the full pipeline: enriching
transactions (categorizing short, noisy merchant strings in mixed English and latinized Arabic,
and resolving them to real merchant entities), then building the models and analyses on top that
turn transaction data into insight across our services - categorization, pattern recognition,
anomaly detection, and risk assessment.
You will be the data scientist for transaction data across the company: owning the enrichment
models end to end (modeling, evaluation, and the production code), and partnering with product
and engineering teams to design, ship, and measure model-backed features in other services.
- Own and level up transaction enrichment - the foundation everything else stands on: categorization that generalizes to merchants we've never seen, and merchant name matching that tolerates harmless variation without confusing genuinely different businesses.
- Own model confidence across our systems: well-calibrated probabilities, principled abstention on uncertain cases, and confidence-based routing that products and review workflows rely on.
- Build transaction-based insight models for other services: pattern recognition, anomaly detection, and behavioral analysis on transaction streams.
- Own the ongoing validation of our risk models: backtesting scores against realized users behavior, monitoring discrimination and calibration as behaviors evolve, and driving model improvements from what the data shows.
- Handle our text and behavior data as they really are: bilingual, informally romanized Arabic with unstable spelling; transaction streams with truncation artifacts, bank quirks, and heavy- tailed distributions.
- Build efficient batch pipelines at hundreds-of-millions-record scale, designed to be re-run routinely as models improve.
- Design within the guardrails of a regulated fintech: data governance, privacy, and cybersecurity requirements shape what data we can use, where workloads can run, and which tools we can adopt - you'll build excellent solutions inside those boundaries, working with those teams rather than around them.
- Build the evaluation discipline these systems deserve: labeled datasets, regression test suites for model behavior, and metrics that answer "did this change help?" for every release.
- 3+ years of applied ML/data science, with models shipped and maintained in production at scale - and knowledge of their failure modes.
- Strong analytical range beyond modeling: exploratory analysis, statistical rigor, and feature design on behavioral/tabular data (SQL fluency assumed).
- Hands-on experience with text similarity, fuzzy matching, or entity resolution on noisy real- world strings.
- Strong Python and the scientific stack (scikit-learn, scipy, numpy), with performance and cost awareness at scale - vectorization, sparse data structures, efficient batch computation.
- Demonstrated maturity working under externally imposed constraints - data governance, privacy, cybersecurity, compliance, infrastructure policy - including collaborating with the teams who own them.
- A track record of turning ambiguous "the model feels wrong" complaints into measured, regression-tested properties.
- Comfort reading and debugging model code you didn't write, and working directly with product teams on loosely-defined problems.
Nice-to-haves
- Arabic / Arabizi text processing - a strong plus; our data is bilingual with unstable romanization.
- Credit or behavioral risk modeling and validation - scorecards, discrimination and calibration measurement, backtesting against realized outcomes.
- Text embeddings and approximate nearest-neighbor retrieval in resource-conscious settings.
- LLM-assisted labeling, distillation, or weak-supervision pipelines.
- ML lifecycle and batch-serving tooling (experiment tracking, data versioning, distributed task queues).