📍 المملكة العربية السعودية تحديث مستمر على مدار الساعة وظائف تناسب سيرتك الذاتيةمجاناً قناة تيليجرام

Fathom.io تعلن عن وظيفة مهندس MLOps في الظهران

MLOps Engineer
🕒 نُشرت: (منذ 18 يوماً) 📍 الظهران وظائف الهندسة والتقنية

تفاصيل الوظيفة

تعلن شركة Fathom.io عن وظيفة MLOps Engineer للعمل في الظهران، المنطقة الشرقية، السعودية. نحن نبحث عن مهندس متوسط إلى كبير للمساهمة في بناء الطبقة الذكية لمنصة الذكاء الاصطناعي الخاصة بنا، حيث ستعمل على تصميم البنية التحتية التي تتيح سير عمل AI متطورة عبر تجربة منتج بسيطة قائمة على الخدمة الذاتية.

المهام والمسؤوليات

  • تصميم وبناء البنية التحتية التي تدعم الطبقة الذكية للمنصة.
  • تمكين سير عمل موثوق وآلي لتدريب النماذج ونشرها وإدارة دورة حياتها والاستدلال.
  • بناء أسس قابلة للتوسع تتيح للمستخدمين إنشاء وتكوين وتشغيل وكلاء AI وسير عمل RAG عبر واجهة المنصة.
  • تطوير إمكانيات المنصة خلف الدفاتر المدارة والوظائف والتجارب ووظائف التدريب وسجلات النماذج ونقاط النهاية للخدمة.
  • تحسين قدرات خدمة النماذج والمراقبة وإدارة الإصدارات والتقييم والترقية والاسترجاع.
  • تحسين عمليات نشر التدريب والاستدلال على وحدات معالجة الرسوميات (GPU) من حيث الأداء والموثوقية وكفاءة التكلفة.
  • استكشاف طرق فعالة لنشر النماذج عبر البنية التحتية المركزية لوحدات معالجة الرسوميات والأجهزة الطرفية.
  • أتمتة سير العمل التي تتطلب جهداً يدوياً في MLOps، مع التركيز على إمكانيات الخدمة الذاتية الآمنة للمستخدمين.
  • التعاون مع فرق الواجهة الخلفية والمنتجات والذكاء الاصطناعي لتحويل البنية التحتية المعقدة إلى ميزات منصة بديهية.
  • المساعدة في وضع معايير للأمان وتعدد المستأجرين وعزل الموارد وحوكمة النماذج والموثوقية التشغيلية.

الشروط والمتطلبات

  • خبرة قوية في MLOps أو هندسة منصات الذكاء الاصطناعي أو البنية التحتية للتعلم الآلي أو الأنظمة الموزعة.
  • خبرة عملية مع Kubernetes، بما في ذلك نشر وتشغيل أعباء العمل الحالة ووظائف التدريب والدفاتر والخوادم بدون خادم أو كثيفة الاستخدام لوحدات معالجة الرسوميات.
  • خبرة في بناء منصات تعلم آلي تدعم مجموعة من التدريب والتجارب والدفاتر وسير عمل البيانات/الميزات وسجلات النماذج والخدمة الإنتاجية.
  • خبرة مع أطر خدمة النماذج مثل KServe أو vLLM أو Triton أو Ray Serve أو ما شابهها.
  • الإلمام بأدوات دورة حياة التعلم الآلي مثل MLflow وتتبع التجارب وسجلات النماذج وCI/CD أو GitOps لأنظمة التعلم الآلي.
  • خبرة عملية في تحسين أعباء عمل التدريب والاستدلال من حيث زمن الاستجابة والإنتاجية والتوافر والتكلفة.
  • فهم جدولة وحدات معالجة الرسوميات وتخصيص الموارد وكمية النماذج والتجميع والتوسع التلقائي وخدمة النماذج المتعددة.
  • خبرة مع أنظمة RAG وتطبيقات LLM ووكلاء الذكاء الاصطناعي أو البنية التحتية الداعمة لها.
  • أسس قوية في هندسة البرمجيات ونهج موجه نحو المنتج في تصميم المنصات.
  • القدرة على جعل سير العمل التشغيلي المعقد موثوقاً وسهل الاستخدام للمستخدمين.

المهارات المطلوبة

  • خبرة في Rust أو اهتمام بالعمل مع أنظمة خلفية مبنية بلغة Rust (ميزة قوية).
  • خبرة في تشغيل بيئات الدفاتر مثل JupyterHub أو Kubeflow Notebooks أو VS Code Server أو ما يعادلها.
  • خبرة في إنشاء بيئات تنفيذ آمنة وقابلة للتطوير للوظائف أو المهام المعرفة من قبل المستخدم.
  • خبرة في نشر أعباء عمل الذكاء الاصطناعي على الأجهزة الطرفية.
  • الإلمام بتقنيات تحسين النماذج، بما في ذلك الكمية والتجميع والتقليم والخدمة المراعية للعتاد.
  • خبرة مع منصات الذكاء الاصطناعي متعددة المستأجرين وضوابط الأمان وحوكمة النماذج.
  • الإلمام بأدوات المراقبة والتقييم لأنظمة التعلم الآلي وLLM.
  • العمل مع تقنيات مثل Kubernetes وKnative وKServe وvLLM وMLflow وLangFuse والبنية التحتية لوحدات معالجة الرسوميات وأعباء خدمة النماذج، وهندسة RAG واسترجاع المتجهات وسير عمل الوكلاء وتطبيقات LLM.
عرض النص الأصلي للإعلان
About The Role

We’re looking for a mid-to-senior MLOps engineer to help build the intelligence layer of our AI platform. This is not a traditional MLOps role focused only on maintaining pipelines and deployments. You will help create the underlying infrastructure that makes sophisticated AI workflows accessible through a simple, self-service product experience: model deployment, training, notebooks, functions, UI-driven agent creation, and RAG pipelines.

You’ll work at the intersection of platform engineering, machine learning infrastructure, distributed systems, and developer experience.

What You’ll Do

  • Design and build the infrastructure powering our Intelligence layer.
  • Enable reliable, automated workflows for model training, deployment, lifecycle management, and inference.
  • Build scalable foundations for users to create, configure, and operate AI agents and RAG pipelines through the platform UI.
  • Develop the platform capabilities behind managed notebooks, functions, experiments, training jobs, model registries, and serving endpoints.
  • Improve model serving, observability, versioning, evaluation, promotion, and rollback capabilities.
  • Optimize GPU inference and training deployments for performance, reliability, and cost efficiency.
  • Explore efficient approaches for deploying models across centralized GPU infrastructure and edge devices.
  • Automate workflows that otherwise require manual MLOps effort, with a focus on safe, self-service capabilities for platform users.
  • Partner with backend, product, and AI teams to turn complex infrastructure into intuitive platform features.
  • Help define standards for security, multi-tenancy, resource isolation, model governance, and operational reliability.

Our current stack

You’ll Work With Technologies Including

  • Kubernetes and Knative
  • KServe and vLLM
  • MLflow
  • LangFuse
  • GPU infrastructure and model-serving workloads
  • RAG architectures, vector retrieval, agent workflows, and LLM applications

Our intelligence backend is built in Rust, so Rust experience is a strong advantage.

What We’re Looking For

  • Strong experience in MLOps, AI platform engineering, machine learning infrastructure, or distributed systems.
  • Hands-on Kubernetes experience, including deploying and operating stateful, training, notebook, serverless, or GPU-intensive workloads.
  • Experience building ML platforms that support some combination of training, experimentation, notebooks, feature/data workflows, model registries, and production serving.
  • Experience with model serving frameworks such as KServe, vLLM, Triton, Ray Serve, or similar.
  • Familiarity with ML lifecycle tooling such as MLflow, experiment tracking, model registries, and CI/CD or GitOps for ML systems.
  • Practical experience optimizing training and inference workloads for latency, throughput, availability, and cost.
  • Understanding of GPU scheduling, resource allocation, model quantization, batching, autoscaling, and multi-model serving.
  • Experience with RAG systems, LLM applications, AI agents, or their supporting infrastructure.
  • Strong software engineering fundamentals and a product-minded approach to platform design.
  • Ability to make complex operational workflows reliable and approachable for users.

Nice to have

  • Rust experience or interest in working with Rust-based backend systems.
  • Experience running notebook environments such as JupyterHub, Kubeflow Notebooks, VS Code Server, or equivalent.
  • Experience creating secure, scalable execution environments for user-defined functions or jobs.
  • Experience deploying AI workloads on edge devices.
  • Familiarity with model optimization techniques, including quantization, compilation, pruning, and hardware-aware serving.
  • Experience with multi-tenant AI platforms, security controls, and model governance.
  • Familiarity with observability and evaluation tooling for ML and LLM systems.

What Success Looks Like

You will help us evolve from manually operated AI infrastructure to a platform where users can train, evaluate, deploy, and operate models; work in managed notebooks; run functions; create agents; and configure RAG workflows - with the platform handling as much of the operational complexity as possible.

By submitting this application, I agree that my personal data will be collected, processed, and retained by the company solely for the purposes of managing and assessing my candidacy.
المصدر: LinkedIn - أُضيفت للموقع في 21 سبتمبر 2026
رقم الإعلان لدى المصدر: 4467384574