وظيفة مهندس أول (Core & MLOps) شاغرة لدى Jobgether في السعودية
تفاصيل الوظيفة
يُعلن موقع Jobgether بالنيابة عن إحدى الشركات الشريكة عن توفر وظيفة Staff Engineer (Core & MLOps) في المملكة العربية السعودية. تتيح هذه الوظيفة فرصة تشكيل البنية التحتية الأساسية التي تدعم منتجات بيانات الويب واسعة النطاق وفرق الهندسة الموزعة.
المهام والمسؤوليات
- تصميم وتطوير مستويي التحكم والسياق (control and context planes) وتطوير سجلات الخدمات والمخططات (service and schema registries) وفرض مؤشرات مستوى الخدمة (SLOs) والتوجيه الصحي والإصدار الآلي الاختباري (canary releases) وحلقات التغذية الراجعة التشغيلية.
- امتلاك حزمة الخدمات (service chassis) والمسار الذهبي (golden path) وصيانة وتحسين مكتبات Java وPython متعددة اللغات ومواصفات أعباء العمل الموحدة وقوالب Helm وخطوط أنابيب النشر.
- تحديد وإدارة العقود البينية للخدمات بما يشمل تعريفات gRPC وProtocol Buffer وترميز API Gateway وسياسات الإصدار ومعايير تطور المخططات.
- تشغيل وتحسين البنية التحتية الأساسية للمنصة عبر Kubernetes وTerraform وHAProxy/Nginx وConfluent Kafka وخطوط أنابيب الفوترة الفورية وValkey ومبادرات تحديث قواعد البيانات.
- قيادة الاستراتيجية المعمارية عبر مناقشات الطلب (RFDs) التي تغطي تنسيق سير العمل وتنسيق البوابة والتوجيه متعدد المجموعات والتبديل التلقائي عند الفشل ومبادرات المنصة الحرجة الأخرى.
- وضع ممارسات هندسة الموثوقية بما فيها مؤشرات مستوى الخدمة (SLIs/SLOs) وميزانيات الأخطاء وعزل الأعطال ونشر الإصدارات الآلي الاختباري المرجح.
- المشاركة في نوبات الدعم للبنية التحتية المشتركة وقيادة تحليلات ما بعد الحوادث وتحويل الرؤى التشغيلية إلى تحسينات في المنصة.
- توجيه المهندسين عبر الفرق المتعددة ومراجعة المقترحات المعمارية ووضع ممارسات هندسية تجعل تطوير البرامج الموثوقة أكثر اتساقاً وكفاءة.
الشروط والمتطلبات
- 10+ سنوات من الخبرة في بناء أنظمة خلفية موزعة قابلة للتوسع مع سجل حافل من إنشاء منصات داخلية أو مكتبات أساسية معتمدة عبر المؤسسات الهندسية.
- خبرة متقدمة في Java بما في ذلك الأطر التفاعلية مثل Vert.x أو Netty بالإضافة إلى إتقان قوي للغة Python.
- خبرة عميقة مع gRPC وProtocol Buffers بما في ذلك تطور المخططات والتوافق العكسي في الأنظمة الحيوية.
- خبرة عملية إنتاجية مع Kubernetes على نطاق واسع وTerraform ومنصات تدفق الأحداث مثل Kafka.
- خبرة في تصميم خطوط أنابيب القياس عن بُعد الآلية (telemetry pipelines) أو طرق العرض المادية (materialized views) أو متاجر الميزات (feature stores) أو غيرها من أنظمة التغذية الراجعة التي تستخدم بيانات الإنتاج لتحسين سلوك النظام ديناميكياً.
- خلفية قوية في هندسة الموثوقية تشمل تعريف SLO/SLI وتحليل نصف قطر الانفجار (blast-radius) وتحمل الأعطال وعقود الخدمة الصارمة.
- مهارات استثنائية في الكتابة الفنية والقدرة على توصيل المفاهيم المعمارية المعقدة بوضوح مع دفع التوافق بين الفرق.
- مهارات تواصل كتابية وشخصية قوية مناسبة لبيئة عمل موزعة عالمياً وتعتمد على العمل عن بُعد.
- عقلية فضولية ورغبة مستمرة في التعلم مع اهتمام بتقييم التقنيات والهندسات الجديدة.
- خبرة مع Temporal أو DBOS أو منصات التنفيذ المتينة المماثلة تعتبر ميزة إضافية.
- خبرة في MLOps تشمل تقديم النماذج ومراقبة الأداء أو كشف الانجراف في الإنتاج تعتبر ميزة.
- الإلمام بالشبكات القائمة على الثقة الصفرية (zero-trust) وشبكات الخدمات مثل SPIRE أو mTLS أو Cilium أو Istio أو Envoy يعتبر مفيداً.
- خبرة في بناء أدوات للمطورين مثل CLIs أو SDKs أو مولدات المشاريع تعتبر ميزة إضافية.
- خبرة في كشط أو زحف الويب على نطاق واسع أو مساهمات في مشاريع مفتوحة المصدر للأنظمة الموزعة واستخراج البيانات تعتبر ميزة.
المزايا
- بيئة عمل مرنة عن بُعد بالكامل مع ساعات عمل مرنة.
- حرية ومرونة في العمل من الموقع الذي يجعلك أكثر إنتاجية.
- فرصة العمل على البنية التحتية الأساسية التي تدعم خطوط أنابيب بيانات الويب واسعة النطاق والأنظمة الموزعة.
- التعرض لأحدث التقنيات والأدوات مفتوحة المصدر وتطور البنية التحتية للويب والذكاء الاصطناعي.
- فرص لحضور المؤتمرات والتواصل مع الزملاء حول العالم.
- التعاون مع مجتمع هندسي متنوع ومتعدد الثقافات وموزع عالمياً.
- درجة عالية من الاستقلالية والثقة التنظيمية.
- فرص للتأثير في هندسة المنصة ومعايير الهندسة والاستراتيجية التقنية عبر فرق متعددة.
عرض النص الأصلي للإعلان
This role offers the opportunity to shape foundational infrastructure powering large-scale web data products and distributed engineering teams. You’ll own the architecture of core control and context planes that enable services and AI-driven workflows to operate reliably and efficiently. Working across Kubernetes, Kafka, Java, Python, gRPC, and multi-cloud infrastructure, you’ll tackle complex distributed systems challenges at production scale. You’ll establish engineering standards, reliability practices, and service contracts that influence multiple product squads. The role combines hands-on architecture with technical leadership, mentoring, and cross-functional alignment. In a globally distributed, remote-first environment, you’ll have significant autonomy to solve challenging infrastructure problems and influence long-term platform strategy.
Accountabilities
- Architect and evolve the control and context planes, advancing service and schema registries, SLO enforcement, health-aware routing, automated canary releases, and operational feedback loops.
- Own the service chassis and golden path, maintaining and improving multi-language Java and Python client libraries, standardized workload specifications, Helm charts, and deployment pipelines.
- Define and govern inter-service contracts, including gRPC and Protocol Buffer definitions, API gateway transcoding, versioning policies, and schema evolution standards.
- Operate and improve the core platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, real-time billing pipelines, Valkey, and database modernization initiatives.
- Lead architectural strategy through Requests for Discussion (RFDs) covering workflow orchestration, gateway orchestration, multi-cluster routing, automated failover, and other critical platform initiatives.
- Establish reliability engineering practices, including SLOs, SLIs, error budgets, fault isolation, and automated weighted canary deployments.
- Participate in shared infrastructure on-call rotations, lead incident post-mortems, and convert operational insights into platform improvements.
- Mentor engineers across multiple squads, review architectural proposals, and establish engineering practices that make reliable software development more consistent and efficient.
- 10+ years of experience building scalable distributed backend systems, with a strong track record of creating internal platforms or core libraries adopted across engineering organizations.
- Advanced Java expertise, including reactive frameworks such as Vert.x or Netty, combined with strong Python proficiency.
- Deep experience with gRPC and Protocol Buffers, including schema evolution and backward compatibility in mission-critical systems.
- Hands-on production experience with Kubernetes at scale, Terraform, and event-streaming platforms such as Kafka.
- Experience designing automated telemetry pipelines, materialized views, feature stores, or other feedback systems that use production data to dynamically improve system behavior.
- Strong reliability engineering background, including SLO/SLI definition, blast-radius analysis, fault tolerance, and rigorous service contracts.
- Exceptional technical writing skills and the ability to communicate complex architectural concepts clearly while driving alignment across teams.
- Strong written and interpersonal communication skills suited to a globally distributed, remote-first environment.
- A curious, continuous-learning mindset with an interest in evaluating new technologies, architectures, and engineering approaches.
- Experience with Temporal, DBOS, or similar durable execution platforms is a plus.
- MLOps experience, including model serving, performance monitoring, or production drift detection, is advantageous.
- Familiarity with zero-trust networking and service meshes such as SPIRE, mTLS, Cilium, Istio, or Envoy is beneficial.
- Experience building developer tooling such as CLIs, SDKs, or project generators is a plus.
- Experience with large-scale web scraping or crawling, or contributions to distributed-systems and data-extraction open-source projects, is advantageous.
- Fully remote, remote-first working environment with flexible working hours.
- Freedom and flexibility to work from the location where you are most productive.
- Opportunity to work on core infrastructure supporting large-scale web data pipelines and distributed systems.
- Exposure to cutting-edge open-source technologies, tools, and evolving AI and web data infrastructure.
- Opportunities to attend conferences and connect with colleagues across the globe.
- Collaboration with a diverse, multicultural, and globally distributed engineering community.
- High level of autonomy and organizational trust.
- Opportunities to influence platform architecture, engineering standards, and technical strategy across multiple teams.
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
رقم الإعلان لدى المصدر: 4472529841