وظيفة مهندس نشر أمامي - MLOps شاغرة لدى Systems Limited
تفاصيل الوظيفة
انضم إلى Systems Limited في السعودية كمهندس نشر أمامي متخصص في MLOps (Forward Deployed Engineer - MLOps)، حيث ستتولى مسؤولية دورة حياة نماذج التعلم الآلي بالكامل، بدءًا من النشر الموثوق والمراقبة والتحسين والصيانة على نطاق واسع.
المهام والمسؤوليات
- إدارة النشر المُنتَج (production serving) وخطوط CI/CD ونشر ومراقبة نماذج التعلم الآلي.
- بناء وصيانة خطوط إعادة تدريب النماذج، وإدارة الإصدارات، ونشرها.
- إدارة البنية التحتية للتعلم الآلي وتحسين تكاليف السحابة عبر ممارسات FinOps.
- تنفيذ المراقبة والتنبيهات، وكشف انحراف النموذج (model drift)، ومراقبة الأداء.
- تحمل مسؤولية الاستجابة للحوادث المُنتَجية (production incidents)، واستكشاف الأخطاء وإصلاحها، والمشاركة في المناوبات (on-call).
- التعاون مع علماء البيانات ومهندسي التعلم الآلي لتصميم أنظمة قابلة للتطوير وجاهزة للإنتاج.
- دعم مرحلة ما قبل البيع (presales) وإثباتات المفهوم (PoCs) من خلال إظهار الجاهزية للإنتاج وقابلية التوسع.
- توجيه المهندسين حول أفضل ممارسات MLOps والجاهزية للإنتاج، والتواصل مع الجهات غير التقنية حول مقايضات التكلفة والأداء والموثوقية.
الشروط والمتطلبات
- 6+ سنوات من الخبرة في MLOps أو هندسة منصات التعلم الآلي أو أدوار مشابهة مع إثبات ملكية الإنتاج.
- خبرة قوية في CI/CD، والحاويات (containerization)، والبنية التحتية السحابية، ومراقبة التعلم الآلي.
- فهم عميق لدورة حياة نماذج التعلم الآلي، بما في ذلك إعادة التدريب، وإدارة الإصدارات، وكشف الانحراف (drift detection).
- خبرة في البنية التحتية كرمز (Infrastructure as Code - IaC) وخطوط النشر الآلية.
- معرفة قوية بمنصات السحابة الرئيسية مثل AWS أو Azure أو GCP.
- خبرة في مراقبة تكاليف السحابة وتحسينها وممارسات FinOps.
- فهم قوي للمراقبة والتنبيهات وإدارة اتفاقيات مستوى الخدمة (SLA) والاستجابة للحوادث المنتَجية.
- القدرة على استكشاف الأخطاء وإصلاحها والتواصل حول الحوادث التقنية بوضوح مع أصحاب المصلحة من قطاع الأعمال.
- مهارات تعاون وتوجيه قوية.
- الرغبة في المشاركة في الدعم الإنتاجي خارج ساعات العمل والمناوبات (on-call).
عرض النص الأصلي للإعلان
Own the production lifecycle of machine learning models, ensuring validated models are reliably deployed, monitored, optimized, and maintained at scale. The role focuses on MLOps, cloud infrastructure, automation, observability, cost optimization, and production incident management.
Responsibilities:
- Own production serving, CI/CD, deployment, and monitoring of ML models.
- Build and maintain model retraining, versioning, and deployment pipelines.
- Manage ML infrastructure and optimize cloud costs through FinOps practices.
- Implement observability, alerting, model drift detection, and performance monitoring.
- Own production incident response, troubleshooting, and on-call responsibilities.
- Collaborate with Data Scientists and ML Engineers to design scalable, production-ready systems.
- Support presales and PoCs by demonstrating production readiness and scalability.
- Mentor engineers on MLOps and production-readiness best practices.
- Communicate infrastructure cost, performance, and reliability trade-offs to non-technical stakeholders.
Qualifications:
- 6+ years of experience in MLOps, ML Platform Engineering, or related roles with proven production ownership.
- Strong expertise in CI/CD, containerization, cloud infrastructure, and ML observability.
- Deep understanding of the ML model lifecycle, including retraining, model versioning, and drift detection.
- Experience with Infrastructure as Code (IaC) and automated deployment pipelines.
- Strong knowledge of major cloud platforms such as AWS, Azure, or GCP.
- Experience with cloud cost monitoring, optimization, and FinOps practices.
- Strong understanding of monitoring, alerting, SLA management, and production incident response.
- Ability to troubleshoot and communicate technical incidents clearly to business stakeholders.
- Strong collaboration and mentoring skills.
- Willingness to participate in on-call and off-hours production support
رقم الإعلان لدى المصدر: 4462799110