تفاصيل الوظيفة
شركة عزم السعودية تعلن عن وظيفة مهندس ذكاء اصطناعي (AI Engineer) في الرياض.
المهام والمسؤوليات
- بناء وضبط وتقييم أنظمة LLM للمهام التخصصية باستخدام QLoRA/PEFT على نماذج مفتوحة مثل Llama-3 وMistral.
- تصميم منصات تقييم قابلة لإعادة الإنتاج وأُطر اختبار A/B مع مقاييس تتبع: معدل نجاح المهمة، معدل الأمان، وتوزيعات زمن الاستجابة (p50/p95).
- هندسة أنظمة متعددة الوكلاء و RAG (LangGraph, FastAPI, قواعد بيانات متجهية) من النموذج الأولي إلى الإنتاج.
- تنفيذ حواجز أمان (guardrails) للتحقق من المدخلات والمخرجات، سياسات القائمة البيضاء/السوداء، وضوابط تقلل الإجراءات غير الصالحة أو عالية المخاطر للنموذج.
- ترجمة حالات الاستخدام التجاري إلى نماذج أولية قابلة للنشر بمعايير قبول قابلة للقياس، وعرضها على أصحاب المصلحة.
- تصميم وتشغيل البنية التحتية السحابية وبيئات MLOps (Azure, OCI, أو GCP) لأعباء العمل الخاصة بالذكاء الاصطناعي على Kubernetes وبيئات التشغيل المحتواة.
- بناء خطوط CI/CD وآليات تعزيز الإصدار المستندة إلى GitOps (Argo CD) عبر بيئات التطوير والاختبار والإنتاج.
- تنفيذ المراقبة الشاملة (Azure Monitor, Application Insights, ELK) مع أهداف كشف واستجابة محددة.
- تطبيق أساسيات أمان الشبكات والمحيط (FW/WAF)، والفحص الآلي لجودة الكود (SonarQube, Black Duck)، وخطوط الأنابيب المقيدة.
- إدارة تصميم التعافي من الكوارث - نسخ احتياطية آلية، تجاوز الفشل، والتزامات RTO/RPO موثقة.
- قيادة وإرشاد فريق عمليات سحابية/ذكاء اصطناعي؛ تحديد ممارسات المراقبة والاستجابة للحوادث وإدارة الإصدار بمعايير واضحة لوقت التشغيل ومتوسط وقت الإصلاح (MTTR).
- توحيد ممارسات دورة حياة تطوير البرمجيات (SDLC): استراتيجية التفرع، حوكمة طلبات السحب، إدارة الإصدار، تقارير التسليم - لتحسين زمن التوصيل وتكرار النشر.
- دمج أدوات الهندسة وسير العمل؛ قيادة عمليات الترحيل وتوحيد المنصات حيث يعيق التجزؤ سرعة التوصيل.
- إنتاج وثائق التسليم وكتيبات التشغيل (runbooks) التي تجعل الأنظمة قابلة للتدقيق وقابلة للنقل تشغيلياً.
- دعم مفاوضات البائعين والتراخيص للاتفاقيات المؤسسية السحابية.
الشروط والمتطلبات
- درجة البكالوريوس في هندسة البرمجيات، علوم الحاسب، أو مجال ذي صلة.
- خبرة من 6 إلى 8+ سنوات في هندسة البرمجيات، DevOps، أو هندسة المنصات، بما في ذلك سنتان على الأقل في مجال هندسة الذكاء الاصطناعي التطبيقي أو تعلم الآلة.
- تقديم أنظمة ذكاء اصطناعي/LLM إنتاجية - وليس فقط أعمالاً بحثية أو على مستوى الملاحظات.
- إتقان قوي للغة Python؛ إلمام بـ Bash وYAML.
- خبرة عملية عميقة مع Kubernetes، Docker/Podman، وTerraform.
- خبرة إنتاجية مع مزود سحابي رئيسي واحد على الأقل (Azure مفضل؛ OCI أو GCP مقبول).
- ملكية مثبتة لأنظمة CI/CD على نطاق واسع (Azure DevOps، GitHub Actions) ونماذج الإصدار المستندة إلى GitOps.
- خبرة في قيادة فريق ووضع معايير هندسية عبر عدة فرق.
- يفضّل: درجة الماجستير في الذكاء الاصطناعي التطبيقي، تعلم الآلة، أو تخصص ذي صلة.
- يفضّل: خبرة في الضبط الدقيق باستخدام QLoRA/LoRA على مجموعات وحدات معالجة رسومية (GPU)؛ إلمام بـ PyTorch وTransformers.
- يفضّل: خبرة في قواعد البيانات المتجهية (Milvus، Pinecone، أو Weaviate) وتصميم استرجاع RAG.
- يفضّل: خبرة في العمل على منصات حكومية سعودية أو منصات رقمية وطنية واسعة النطاق، مع الإلمام بالامتثال والمعايير المحلية.
- يفضّل: إتقان اللغتين العربية والإنجليزية بمستوى مهني.
- البيئة التقنية: Python، FastAPI، PyTorch، Transformers، LangGraph، Milvus/Pinecone/Weaviate، Redis، PostgreSQL، Kubernetes، Docker/Podman، Terraform، Argo CD، Azure DevOps، GitHub Actions، Azure ML، Azure Monitor / Application Insights، ELK، SonarQube، Black Duck، Fortinet FW/WAF.
عرض النص الأصلي للإعلان
Description
AI systems
Build, fine-tune, and evaluate LLM systems for domain-specific tasks (QLoRA / PEFT on open-weight models such as Llama-3 and Mistral).
Design reproducible evaluation harnesses and A/B test frameworks with tracked metrics: task success rate, safety rate, and latency distributions (p50/p95).
Architect multi-agent and RAG systems (LangGraph, FastAPI, vector databases) from prototype through production.
Implement safety guardrails - input/output validation, allowlist/denylist policies, and controls that reduce invalid or high-risk model actions.
Translate business use cases into deployable prototypes with measurable acceptance criteria, and demo them to stakeholders.
Platform & infrastructure
Design and operate cloud infrastructure and MLOps workspaces (Azure, OCI, or GCP) for AI workloads on Kubernetes and containerized runtimes.
Build CI/CD pipelines and GitOps-based release promotion (Argo CD) across development, test, and production environments.
Implement end-to-end observability (Azure Monitor, Application Insights, ELK) with defined detection and response targets.
Apply network and perimeter security baselines (FW/WAF), automated code quality and SCA scanning (SonarQube, Black Duck), and gated pipelines.
Own disaster recovery design - automated backups, failover, and documented RTO/RPO commitments.
Engineering leadership
Lead and mentor a cloud/AI operations team; define monitoring, incident response, and release governance practices with clear uptime and MTTR targets.
Standardize SDLC practices - branching strategy, PR governance, release management, delivery reporting - to improve lead time and deployment frequency.
Consolidate engineering tooling and workflows; drive migrations and platform standardization where fragmentation slows delivery.
Produce handover documentation and runbooks that make systems auditable and operationally transferable.
Support vendor and licensing negotiations for cloud enterprise agreements.
Requirements
Required Qualifications
Bachelor's degree in Software Engineering, Computer Science, or a related field.
6-8+ years in software, DevOps, or platform engineering, including at least 2 years in an applied AI or ML engineering capacity.
Proven delivery of production AI/LLM systems - not only research or notebook-stage work.
Strong Python; comfortable with Bash and YAML.
Deep hands-on experience with Kubernetes, Docker/Podman, and Terraform.
Production experience with at least one major cloud (Azure preferred; OCI or GCP acceptable).
Demonstrated ownership of CI/CD at scale (Azure DevOps, GitHub Actions) and GitOps release models.
Experience leading a team and setting engineering standards across multiple squads.
Preferred Qualifications
Master's degree in Applied AI, Machine Learning, or a related discipline.
Fine-tuning experience with QLoRA/LoRA on GPU clusters; PyTorch and Transformers.
Vector database experience (Milvus, Pinecone, or Weaviate) and RAG retrieval design.
Experience delivering on Saudi government or large-scale national digital platforms, with familiarity in local compliance and standards.
Arabic and English professional proficiency.
Technical Environment
Python, FastAPI, PyTorch, Transformers, LangGraph, Milvus/Pinecone/Weaviate, Redis, PostgreSQL, Kubernetes, Docker/Podman, Terraform, Argo CD, Azure DevOps, GitHub Actions, Azure ML, Azure Monitor / Application Insights, ELK, SonarQube, Black Duck, Fortinet FW/WAF.
رقم الإعلان لدى المصدر: 2716855