📍 المملكة العربية السعودية تحديث مستمر على مدار الساعة وظائف تناسب سيرتك الذاتيةمجاناً قناة تيليجرام

وظيفة مهندس أبحاث ذكاء اصطناعي (تحسين النواة والاستدلال) لدى Jobgether في السعودية

AI Research Engineer (Kernel & Inference Optimization)
🕒 نُشرت: (منذ 9 أيام) 📍 السعودية وظائف الهندسة والتقنية

تفاصيل الوظيفة

Jobgether تعلن، نيابة عن شركة شريكة، عن توفر وظيفة مهندس أبحاث ذكاء اصطناعي (تحسين النواة والاستدلال) في السعودية. ستعمل على تقاطع أبحاث الذكاء الاصطناعي وهندسة الأنظمة والاستدلال عالي الأداء، مع التركيز على تطوير بنى خدمة النماذج وتحسينها لمجموعة من البيئات الحاسوبية، بما في ذلك الأجهزة المحمولة والحافة.

المهام والمسؤوليات

  • تصميم ونشر بنى خدمة نماذج متقدمة مُحسَّنة لتحقيق إنتاجية عالية وزمن استجابة منخفض واستخدام فعال للذاكرة.
  • تطوير خطوط أنابيب استدلال قادرة على العمل بفعالية عبر بيئات متنوعة، بما في ذلك الأجهزة المحمولة ومنصات الحافة محدودة الموارد.
  • وضع أهداف أداء واضحة تشمل زمن استجابة الاستجابة، سرعة توليد الرموز، الإنتاجية، بصمة الذاكرة، والموثوقية.
  • بناء وتنفيذ مقاييس أداء استدلال محكومة في بيئات محاكاة وإنتاج، مع تتبع زمن الاستجابة، الإنتاجية، استهلاك الذاكرة، ومعدلات الأخطاء.
  • إنشاء وصيانة مجموعات بيانات وسيناريوهات محاكاة تمثيلية لتقييم أداء النموذج تحت ظروف حقيقية ومحدودة الموارد.
  • تحديد الاختناقات الحسابية والذاكرية عبر خطوط أنابيب الاستدلال وتنفيذ حلول تتضمن التجميع، الشبكات، إدارة الذاكرة، وتحسينات على مستوى النظام.
  • تطوير نوى GPU مخصصة وشادرات حاسوبية للأجهزة المحمولة، بما في ذلك حلول مكتوبة بلغة Metal Shading Language (MSL).
  • تطبيق تقنيات تحسين استدلال متقدمة مثل التشذيب (pruning)، التكميم (quantization)، Flash Attention، التخزين المؤقت KV (KV caching)، وفك التشفير التخميني (speculative decoding).
  • تصميم وتحسين أنظمة الاستدلال الموزعة باستخدام أساليب مثل توازي الموتر (tensor parallelism)، توازي الأنابيب (pipeline parallelism)، وتوازي الخبراء (expert parallelism) لأحمال عمل GPU واسعة النطاق.
  • العمل مع فرق هندسة وأبحاث متعددة التخصصات لدمج أطر الاستدلال المُحسَّنة في التطبيقات الإنتاجية وتطبيقات الأجهزة الطرفية.
  • تحديد منهجيات التقييم، توثيق النتائج التجريبية، مقارنة الأداء مع المعايير المعمول بها، والتحسين المستمر لاستراتيجيات التحسين.
  • مراقبة أداء الإنتاج واستخدام الأبحاث التجريبية لتحديد فرص التحسين في قابلية التوسع والكفاءة والموثوقية.

الشروط والمتطلبات

  • درجة في علوم الحاسب أو مجال تقني ذي صلة؛ يُعتبر الدكتوراه في معالجة اللغات الطبيعية (NLP) أو التعلم الآلي أو تخصص ذي صلة أمراً بالغ الأهمية، خاصة مع سجل أبحاث قوي في الذكاء الاصطناعي ومنشورات في مؤتمرات رائدة.
  • خبرة مثبتة في لغة Metal Shading Language (MSL)، بما في ذلك القدرة على كتابة شادرات حاسوبية مخصصة من الصفر.
  • خبرة موثقة في تحسين النواة منخفض المستوى (low-level kernel optimization) وتحسين الاستدلال على الأجهزة المحمولة أو غيرها من الأجهزة محدودة الموارد.
  • سجل حافل في تحقيق تحسينات قابلة للقياس في زمن استجابة الاستدلال، الإنتاجية، وبصمة الذاكرة لتطبيقات مجال محددة.
  • فهم عميق لبنى خدمة النماذج الحديثة، محركات الاستدلال، وتقنيات التحسين لنشر الذكاء الاصطناعي عالي الأداء.
  • خبرة قوية في كتابة نوى GPU للأجهزة المحمولة مثل الهواتف الذكية.
  • خبرة عملية في تطوير ونشر خطوط أنابيب استدلال كاملة - من تحسين النموذج إلى التكامل الإنتاجي على أجهزة محدودة الموارد.
  • قدرة قوية على تطبيق الأبحاث التجريبية والتجارب المنهجية لحل تحديات زمن الاستجابة والحوسبة والذاكرة.
  • خبرة في تصميم أطر تقييم وقياس أداء قوية لأنظمة الاستدلال.
  • معرفة بتقنيات الاستدلال الموزع، بما في ذلك توازي الموتر، توازي الأنابيب، وتوازي الخبراء لمجموعات GPU واسعة النطاق.
  • فهم عميق للأسس الرياضية وهندسة نماذج الانتشار (diffusion models) ومحولات الرؤية (Vision Transformers - ViTs).
  • إلمام بتقنيات تحسين الاستدلال الحديثة مثل التشذيب، التكميم، Flash Attention، تحسين KV Cache، وفك التشفير التخميني مثل EAGLE.
  • مهارات تحليلية وحل مشكلات قوية، مع القدرة على التحقيق في اختناقات النظام المعقدة وتحويل نتائج الأبحاث إلى حلول هندسية عملية.
  • مهارات اتصال ممتازة باللغة الإنجليزية والقدرة على التعاون بفعالية مع فرق تقنية موزعة ومتعددة التخصصات.

المزايا

  • فرصة العمل على أنظمة ذكاء اصطناعي متقدمة تشمل خدمة النماذج، تحسين الاستدلال، الحوسبة المحمولة، النشر على الحافة، والاستدلال الموزع واسع النطاق.
  • بيئة عمل تعتمد العمل عن بُعد أولاً (remote-first) مع فريق دولي.
  • التعرض لأبحاث الذكاء الاصطناعي المتطورة وتحديات الهندسة العملية.
  • فرصة المساهمة في بنية تحتية حساسة للأداء حيث يمكن للتحسينات أن تحقق أثراً قابلاً للقياس على تطبيقات الذكاء الاصطناعي الحقيقية.
  • بيئة تعاونية تجمع بين التجارب البحثية والهندسة العملية.
  • فرصة العمل مع بنى نماذج متقدمة تشمل نماذج الانتشار ومحولات الرؤية والأنظمة متعددة الوسائط.
  • الوصول إلى مشاكل تقنية صعبة تتضمن نوى GPU ومحركات الاستدلال وتحسين الذاكرة والحوسبة الموزعة.
عرض النص الأصلي للإعلان
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in Saudi Arabia.

You will work at the intersection of AI research, systems engineering, and high-performance model inference.

Your focus will be on developing and optimizing model-serving architectures for advanced AI systems across a range of hardware environments.

You will tackle challenges involving latency, throughput, memory efficiency, and scalability, including deployment on resource-constrained mobile and edge devices.

The role combines hands-on research with low-level engineering, giving you the opportunity to develop novel inference strategies and GPU kernels.

You will work with complex architectures spanning text, image, audio, diffusion models, and vision transformers.

Your work will involve rigorous benchmarking, production testing, and iterative optimization to translate research into measurable performance improvements.

You will collaborate with cross-functional teams in a highly technical, remote environment focused on pushing the boundaries of efficient AI systems.

Accountabilities

  • Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.
  • Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.
  • Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.
  • Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.
  • Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.
  • Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations.
  • Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL).
  • Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.
  • Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads.
  • Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications.
  • Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies.
  • Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability.

Requirements

  • Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences.
  • Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch.
  • Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices.
  • Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications.
  • Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment.
  • Strong experience writing GPU kernels for mobile devices such as smartphones.
  • Practical experience developing and deploying end-to-end inference pipelines, from model optimization through production integration on constrained hardware.
  • Strong ability to apply empirical research and systematic experimentation to solve latency, computational, and memory challenges.
  • Experience designing robust evaluation and benchmarking frameworks for inference systems.
  • Knowledge of distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU clusters.
  • Deep understanding of the mathematical foundations and architecture of diffusion models and Vision Transformers.
  • Familiarity with modern inference optimization techniques including pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding such as EAGLE.
  • Strong analytical and problem-solving abilities, with an ability to investigate complex system bottlenecks and turn research findings into practical engineering solutions.
  • Excellent English communication skills and the ability to collaborate effectively with distributed, cross-functional technical teams.

Benefits

  • Opportunity to work on advanced AI systems spanning model serving, inference optimization, mobile computing, edge deployment, and large-scale distributed inference.
  • Remote-first working environment with an international team.
  • Exposure to cutting-edge AI research and practical systems engineering challenges.
  • Opportunity to contribute to performance-critical infrastructure where improvements can have a measurable impact on real-world AI applications.
  • Collaborative environment combining research-driven experimentation with hands-on engineering.
  • Opportunity to work with advanced model architectures including diffusion models, Vision Transformers, and multimodal systems.
  • Access to challenging technical problems involving GPU kernels, inference engines, memory optimization, and distributed computing.

How Jobgether Works

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

المصدر: LinkedIn - أُضيفت للموقع في 30 سبتمبر 2026
رقم الإعلان لدى المصدر: 4472151508