تفاصيل الوظيفة
تسعى Devsinc إلى تعيين مهندس بيانات أول (Senior Data Engineer) ذو خبرة من 4 إلى 6 سنوات في مجال تصميم وبناء وصيانة خطوط أنابيب بيانات قابلة للتوسع، وذلك للعمل في الرياض، السعودية.
المهام والمسؤوليات
- تصميم وتطوير وصيانة خطوط أنابيب ETL/ELT قابلة للتوسع لاستيعاب البيانات وتحويلها والتحقق من صحتها وتسليمها.
- بناء وجدولة وتنسيق سير عمل بيانات إنتاجي باستخدام Apache Airflow.
- تطوير مهام معالجة بيانات موزعة باستخدام Apache Spark.
- كتابة كود Python فعال وقابل لإعادة الاستخدام وقابل للصيانة لمعالجة البيانات والأتمتة وتطوير خطوط الأنابيب.
- تطوير وتحسين استعلامات SQL المعقدة لتحويل البيانات وتحليلها والتحقق من جودتها.
- تصميم خطوط أنابيب قادرة على معالجة مجموعات بيانات ضخمة منظمة وشبه منظمة وغير منظمة.
- تنفيذ Redis للتخزين المؤقت والوصول عالي الأداء للبيانات وتطبيقات كثيفة البيانات.
- دمج البيانات من واجهات برمجة التطبيقات (APIs) وقواعد البيانات العلائقية والملفات ومزودين خارجيين ومصادر داخلية وخارجية أخرى.
- تنفيذ التحقق من صحة البيانات والمراقبة والتسجيل ومعالجة الأخطاء والتنبيهات وقابلية مراقبة خطوط الأنابيب.
- تحسين أداء خطوط الأنابيب وتخزين البيانات ووقت المعالجة وتكاليف البنية التحتية.
- تطوير أطر عمل قابلة لإعادة الاستخدام لاستيعاب البيانات وتحويلها بدلاً من كتابة سكربتات لمرة واحدة.
- استكشاف أعطال خطوط الأنابيب واختناقات الأداء ومشاكل جودة البيانات وإصلاحها لضمان الحل في الوقت المناسب.
- التعاون مع فرق ذكاء الأعمال والمنتج وعلوم البيانات والهندسة لتقديم مجموعات بيانات موثوقة جاهزة للإنتاج.
- وضع معايير هندسة البيانات والتوثيق الفني وأفضل ممارسات التطوير والحفاظ عليها.
الشروط والمتطلبات
- درجة البكالوريوس في علوم الحاسب أو هندسة البرمجيات أو علوم البيانات أو مجال ذي صلة.
- 4-6 سنوات من الخبرة المهنية في هندسة البيانات أو دور وثيق الصلة.
- خبرة عملية قوية في تصميم وتنفيذ خطوط أنابيب ETL/ELT إنتاجية.
- إتقان قوي للغة Python لمعالجة البيانات والأتمتة وسير عمل هندسة البيانات.
- مهارات متقدمة في SQL وفهم قوي لقواعد البيانات العلائقية.
- خبرة عملية إنتاجية مع Apache Airflow لتنسيق سير العمل.
- خبرة عملية مع Apache Spark ومعالجة البيانات الموزعة.
- فهم قوي لنمذجة البيانات وأنماط التحويل وهندسة خطوط أنابيب البيانات.
- خبرة مع Redis واستراتيجيات التخزين المؤقت وأنماط الوصول عالية الأداء للبيانات.
- خبرة في معالجة مجموعات البيانات الضخمة وتحسين أداء خطوط الأنابيب والاستعلامات.
- فهم قوي لجودة البيانات والتحقق والمراقبة وقابلية المراقبة وموثوقية خطوط الأنابيب.
- ألفة مع Linux وGit وDocker/الحاويات وممارسات هندسة البرمجيات الحديثة.
- مهارات تحليلية قوية واستكشاف الأخطاء وإصلاحها والتواصل والتعاون عبر الفرق.
المهارات المطلوبة (مفضلة)
- خبرة مع منصات البيانات السحابية وخدمات تخزين الكائنات مثل AWS أو Azure أو GCP.
- خبرة في العمل مع PostgreSQL أو مستودعات البيانات أو قواعد البيانات التحليلية.
- خبرة في معالجة البيانات الجغرافية المكانية أو مجموعات البيانات واسعة النطاق القائمة على الموقع.
- ألفة مع DuckDB أو Apache Sedona أو Trino أو Presto أو تقنيات تحليلية مماثلة.
- خبرة في معالجة الأحداث عالية الحجم أو بيانات التنقل أو المعاملات أو البيانات الجغرافية المكانية.
- ألفة مع خطوط أنابيب CI/CD وممارسات البنية التحتية كرمز.
- خبرة في دعم منتجات البيانات أو منصات التحليلات أو خطوط أنابيب AI/ML.
عرض النص الأصلي للإعلان
Devsinc is looking for a highly skilled Senior Data Engineer with 4-6 years of professional experience to design, build, and maintain scalable data pipelines and processing systems that support analytics, data products, and AI/ML capabilities.
The ideal candidate will have strong hands-on experience with ETL/ELT pipelines, Python, SQL, Apache Airflow, Apache Spark, and Redis, along with the ability to develop reliable and maintainable workflows for large, complex, and continuously growing datasets. You will collaborate with Product, Business Intelligence, Data Science, and Engineering teams to deliver production-ready datasets and high-performance data solutions.
Responsibilities
The ideal candidate will have strong hands-on experience with ETL/ELT pipelines, Python, SQL, Apache Airflow, Apache Spark, and Redis, along with the ability to develop reliable and maintainable workflows for large, complex, and continuously growing datasets. You will collaborate with Product, Business Intelligence, Data Science, and Engineering teams to deliver production-ready datasets and high-performance data solutions.
Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines for data ingestion, transformation, validation, and delivery.
- Build, schedule, and orchestrate production-grade data workflows using Apache Airflow.
- Develop distributed data-processing jobs using Apache Spark.
- Write efficient, reusable, and maintainable Python code for data processing, automation, and pipeline development.
- Develop and optimize complex SQL queries for data transformation, analysis, and quality validation.
- Design pipelines capable of processing large-scale structured, semi-structured, and unstructured datasets.
- Implement Redis for caching, high-performance data access, and data-intensive application requirements.
- Integrate data from APIs, relational databases, files, third-party providers, and other internal and external sources.
- Implement data validation, monitoring, logging, error handling, alerting, and pipeline observability.
- Optimize pipeline performance, data storage, processing time, and infrastructure costs.
- Develop reusable data-ingestion and transformation frameworks instead of one-off scripts.
- Troubleshoot pipeline failures, performance bottlenecks, and data-quality issues to ensure timely resolution.
- Collaborate with BI, Product, Data Science, and Engineering teams to deliver reliable, production-ready datasets.
- Establish and maintain data-engineering standards, technical documentation, and development best practices.
- Bachelor's degree in Computer Science, Software Engineering, Data Science, or a related field.
- 4-6 years of professional experience in Data Engineering or a closely related role.
- Strong hands-on experience designing and implementing production-grade ETL/ELT pipelines.
- Strong proficiency in Python for data processing, automation, and data-engineering workflows.
- Advanced SQL skills and a strong understanding of relational databases.
- Practical production experience with Apache Airflow for workflow orchestration.
- Hands-on experience with Apache Spark and distributed data processing.
- Strong understanding of data modelling, transformation patterns, and data-pipeline architecture.
- Experience with Redis, caching strategies, and high-performance data-access patterns.
- Experience processing large datasets and optimizing pipeline and query performance.
- Strong understanding of data quality, validation, monitoring, observability, and pipeline reliability.
- Familiarity with Linux, Git, Docker/containers, and modern software-engineering practices.
- Strong analytical, troubleshooting, communication, and cross-functional collaboration skills.
- Experience with cloud data platforms and object storage services such as AWS, Azure, or GCP.
- Experience working with PostgreSQL, data warehouses, or analytical databases.
- Experience processing geospatial data or large-scale location-based datasets.
- Familiarity with DuckDB, Apache Sedona, Trino, Presto, or similar analytical technologies.
- Experience processing high-volume event, mobility, transactional, or geospatial data.
- Familiarity with CI/CD pipelines and infrastructure-as-code practices.
- Experience supporting data products, analytics platforms, or AI/ML pipelines
المصدر: LinkedIn - أُضيفت للموقع في 24 سبتمبر 2026
رقم الإعلان لدى المصدر: 4471469662
رقم الإعلان لدى المصدر: 4471469662