Qualcomm تعلن عن وظيفة مهندس أول للبنية التحتية وموثوقية المواقع - هندسة مراكز البيانات والذكاء الاصطناعي في الرياض
Senior Staff Infrastructure & Site Reliability Engineer - Datacentre AI Engineering - Riyadh, KSA
🏢 Qualcomm
تفاصيل الوظيفة
شركة Qualcomm تعلن عن وظيفة Senior Staff Infrastructure & Site Reliability Engineer - Datacentre AI Engineering في الرياض، المملكة العربية السعودية، لدعم توسعة البنية التحتية للحوسبة والذكاء الاصطناعي ضمن رؤية 2030.
نبذة عن الوظيفة
- تركيز الدور على تصميم وتشغيل وتحسين أنظمة استدلال الذكاء الاصطناعي (AI inference) واسعة النطاق في بيئة مراكز البيانات.
- المسؤولية عن ضمان موثوقية وقابلية التوسع والجاهزية الإنتاجية لبنية Qualcomm التحتية للذكاء الاصطناعي لدعم أعباء عمل التعلم الآلي المتقدمة.
- يتطلب الدور أساسيات قوية في هندسة النظم والبرمجيات، وقدرة على التنفيذ العملي والعمل المستقل في مجالات معقدة بالتعاون مع فرق متعددة التخصصات (الأجهزة والبرمجيات والتعلم الآلي).
- المرشح المثالي يمتلك خبرة 12+ سنة في مجال SRE.
المهام والمسؤوليات
- تصميم ونشر وتشغيل أنظمة استدلال الذكاء الاصطناعي واسعة النطاق لدعم أعباء العمل الحرجة في AI.
- ضمان موثوقية وتوفر وقابلية التوسع لمجموعات (clusters) الذكاء الاصطناعي في مراكز بيانات Qualcomm.
- تطوير وصيانة أدوات برمجية وبنية تحتية داعمة لمكدسات (stacks) برمجيات AI.
- تحليل متطلبات البرمجيات والتعاون مع مهندسي البنية والأجهزة لدعم أعباء عمل AI.
- بناء ونشر وتشغيل المكونات الداعمة لاستدلال نماذج اللغة الكبيرة (LLM inference) وسير عمل الذكاء الاصطناعي العاملي (agentic AI) وخدمات AI.
- العمل مع فرق النماذج والأنظمة والبرمجيات لتحسين أداء النماذج على نشرات AI100.
- تحديد وتنفيذ التحسينات لأعباء العمل التي تعمل على أنظمة متعددة المعالجات (multi-SoC) ومتعددة البطاقات (multi-card).
- تطبيق أساسيات SRE بما في ذلك المراقبة والتنبيه والاستجابة للحوادث وتحسين الأداء.
- دعم أنظمة ML الإنتاجية باستخدام أدوات MLOps وأفضل الممارسات التشغيلية.
- المساهمة في مراجعات الحوادث، والتوثيق التشغيلي، والتحسين المستمر للموثوقية.
- بناء وصيانة أدوات المراقبة (observability) ولوحات البيانات والتنبيهات لمراقبة صحة النظام وموثوقيته.
- مراقبة البنية التحتية والخدمات باستخدام أدوات مثل Prometheus و Grafana و CloudWatch والتليمتري المخصصة.
- إنشاء وصيانة التوثيق الفني ودفاتر التشغيل (runbooks) ومقالات قاعدة المعرفة.
- تطوير الأتمتة لتقليل المهام التشغيلية اليدوية وتحسين موثوقية النظام.
- دعم خطوط أنابيب CI/CD لنشر خدمات AI والعوامل الذكية (agents).
- تطبيق ممارسات البنية التحتية كرمز (Infrastructure-as-Code) باستخدام أدوات مثل Terraform و Ansible.
المهارات المطلوبة
- الذكاء الاصطناعي والتعلم العميق: خبرة في أعباء عمل AI/ML مثل LLMs و NLP و Vision و Audio أو أنظمة التوصية. فهم مفاهيم استدلال ML (batching, token streaming, الأداء). خبرة عملية مع PyTorch ومعرفة بأطر ML الحديثة. الإلمام بالاستدلال الموزع (distributed inference) والتدقيق (checkpointing) وبيئات الحوسبة المعجّلة (accelerator-based).
- عمليات الذكاء الاصطناعي: خبرة في دعم تطبيقات AI أو ML في بيئات إنتاجية. الإلمام بخطوط أنابيب استدلال LLM وعمليات خدمات AI.
- البرمجة وتصميم البرمجيات: مهارات قوية في Python مع خبرة في بناء ودعم أنظمة إنتاجية. خبرة في كتابة السكربتات والأتمتة باستخدام Python و Bash. الإلمام بأدوات إدارة التكوين والتنسيق (orchestration).
- النظم والبنية التحتية: أساسيات قوية في Linux (الشل، الحاويات، خدمات النظام، أساسيات الشبكات مثل DNS و TLS و HTTP/gRPC). خبرة في العمل مع مجدولات الكتلة مثل Slurm أو أنظمة مكافئة. خبرة في تشغيل الأنظمة الموزعة مع توفر عالٍ وتحمل الأخطاء.
- المراقبة والرصد: خبرة عملية مع أدوات المراقبة والتسجيل مثل Prometheus و Grafana و ELK أو Loki. فهم إدارة الحوادث ومقاييس صحة الخدمة ومراقبة موثوقية النظام.
- ممارسات DevOps و SRE: فهم متين لدورة حياة تطوير البرمجيات (SDLC) وعمليات الإصدار وممارسات الموثوقية التشغيلية. الإلمام بخطوط أنابيب CI/CD وأدوات البنية التحتية كرمز.
- المهارات المفضلة: خبرة مع GenAI أو Agentic AI أو أطر تنسيق LLM. الإلمام بـ LangChain أو AutoGen أو أنظمة RAG. خبرة مع أطر ML إضافية مثل TensorFlow أو JAX أو Ray. معرفة بأنظمة GPU/المسرعات والشبكات عالية الأداء (RDMA, InfiniBand, RoCE). خبرة مع سير عمل MLOps المتقدمة أو عمليات منصات AI واسعة النطاق.
الشروط والمتطلبات
- درجة البكالوريوس أو الماجستير في الهندسة أو علوم الحاسب أو الذكاء الاصطناعي/التعلم الآلي أو مجال ذي صلة.
- 12+ سنة خبرة في هندسة البرمجيات أو النظم أو البنية التحتية، ويفضل أن تكون في بيئات إنتاجية أو مراكز بيانات.
- بكالوريوس في الهندسة أو نظم المعلومات أو علوم الحاسب أو مجال ذي صلة مع 6+ سنوات خبرة في هندسة اختبار البرمجيات أو مجال ذي صلة. أو ماجستير مع 5+ سنوات. أو دكتوراه مع 4+ سنوات.
- سنتان أو أكثر من الخبرة العملية في اختبار البرمجيات أو اختبار النظام وتطوير وأتمتة خطط الاختبار والأدوات (مثل أنظمة التحكم بالمصدر وأدوات التكامل المستمر وأدوات تتبع الأخطاء).
المزايا
- راتب يشمل بدل سكن ومواصلات.
- أسهم (RSUs) ومكافأة مرتبطة بالأداء.
- إجازة أمومة مدفوعة بالكامل لمدة 16 أسبوعًا.
- إجازة أبوة مدفوعة بالكامل لمدة 6 أسابيع.
- خطط شراء أسهم للموظفين.
- بدل تعليم للأطفال.
- دعم الانتقال والتأشيرات (إذا لزم الأمر).
- تأمين على الحياة وتأمين طبي.
- تعويض Live+ Well للاشتراكات الصحية والترفيهية.
عرض النص الأصلي للإعلان
Company
Qualcomm Middle East Information Technology Company LLC
Job Area
Engineering Group, Engineering Group > Software Test Engineering
General Summary
About Us
Qualcomm is growing its presence in Riyadh and is hiring Data Centre Engineers to support our expanding infrastructure across the region. As Saudi Arabia accelerates its digital transformation under Vision 2030, Qualcomm is investing in world‑class computing and data centre capabilities to power AI, cloud, and advanced connectivity at scale. This is a unique opportunity to work in a fast‑growing technology hub, supporting critical environments and helping shape the future of data centre operations in the Kingdom and beyond.
About The Role
The role focuses on the design, operation, and continuous improvement of large‑scale AI inference systems in a datacenter environment. The engineer will support critical AI use cases by ensuring Qualcomm’s AI infrastructure is reliable, scalable, and production‑ready for advanced machine‑learning workloads.
The role requires strong systems and software engineering fundamentals, hands‑on execution, and the ability to work independently on complex problem areas while collaborating closely with cross‑functional teams across hardware, software, and machine learning.
Ideal candidate will have 12+ years of experience in SRE
Key Responsibilities
Key Responsibilities will include
Apart from working with great people, we offer the below:
Master's degree in Engineering, Information Systems, Computer Science, or related field and 5+ years of Software Test Engineering or related work experience.
OR
PhD in Engineering, Information Systems, Computer Science, or related field and 4+ years of Software Test Engineering or related work experience.
Qualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.
To all Staffing and Recruiting Agencies: Our Careers Site is only for individuals seeking a job at Qualcomm. Staffing and recruiting agencies and individuals being represented by an agency are not authorized to use this site or to submit profiles, applications or resumes, and any such submissions will be considered unsolicited. Qualcomm does not accept unsolicited resumes or applications from agencies. Please do not forward resumes to our jobs alias, Qualcomm employees or any other company location. Qualcomm is not responsible for any fees related to unsolicited resumes/applications.
If you would like more information about this role, please contact Qualcomm Careers.
Qualcomm Middle East Information Technology Company LLC
Job Area
Engineering Group, Engineering Group > Software Test Engineering
General Summary
About Us
Qualcomm is growing its presence in Riyadh and is hiring Data Centre Engineers to support our expanding infrastructure across the region. As Saudi Arabia accelerates its digital transformation under Vision 2030, Qualcomm is investing in world‑class computing and data centre capabilities to power AI, cloud, and advanced connectivity at scale. This is a unique opportunity to work in a fast‑growing technology hub, supporting critical environments and helping shape the future of data centre operations in the Kingdom and beyond.
About The Role
The role focuses on the design, operation, and continuous improvement of large‑scale AI inference systems in a datacenter environment. The engineer will support critical AI use cases by ensuring Qualcomm’s AI infrastructure is reliable, scalable, and production‑ready for advanced machine‑learning workloads.
The role requires strong systems and software engineering fundamentals, hands‑on execution, and the ability to work independently on complex problem areas while collaborating closely with cross‑functional teams across hardware, software, and machine learning.
Ideal candidate will have 12+ years of experience in SRE
Key Responsibilities
Key Responsibilities will include
- AI Infrastructure
- Design, deploy, and operate large‑scale AI inference systems supporting critical AI workloads.
- Ensure reliability, availability, and scalability of Qualcomm datacenter AI clusters.
- Develop and maintain software tools and support infrastructure around AI software stacks.
- AI & ML Engineering
- Analyze software requirements and collaborate with architecture and hardware engineers to support AI workloads.
- Build, deploy, and operate components supporting LLM inference, agentic AI workflows, and AI services.
- Work with models, systems, and software teams to improve model performance on AI100 deployments.
- Identify and implement optimizations for workloads running on multi‑SoC and multi‑card systems.
- Site Reliability Engineering (SRE)
- Apply SRE fundamentals including monitoring, alerting, incident response, and performance optimization.
- Support production ML systems using MLOps tools and operational best practices.
- Contribute to incident reviews, operational documentation, and continuous reliability improvements.
- Observability & Tooling
- Build and maintain observability tools, dashboards, and alerts to monitor system health and reliability.
- Monitor infrastructure and services using tools such as Prometheus, Grafana, CloudWatch, and custom telemetry.
- Create and maintain technical documentation, runbooks, and knowledge‑base articles.
- Automation & CI/CD
- Develop automation to reduce manual operational tasks and improve system reliability.
- Support CI/CD pipelines for AI service and agent deployment.
- Apply Infrastructure‑as‑Code practices using tools such as Terraform and Ansible.
- AI & Deep Learning
- Experience working with AI/ML workloads such as LLMs, NLP, Vision, Audio, or Recommendation systems.
- Understand ML inference concepts including batching, token streaming, and performance considerations.
- Hands‑on experience with PyTorch and familiarity with modern ML frameworks.
- Familiarity with distributed inference, checkpointing, and accelerator‑based compute environments.
- AI Operations
- Experience supporting AI or ML applications in production environments.
- Familiarity with LLM inference pipelines and AI service operations.
- Programming & Software Design
- Strong programming skills in Python with experience building and supporting production systems.
- Experience with scripting and automation using Python and Bash.
- Familiarity with configuration management and orchestration tools.
- Systems & Infrastructure
- Strong Linux fundamentals include shell, containers, system services, and networking basics (DNS, TLS, HTTP/gRPC).
- Experience working with cluster schedulers such as Slurm or equivalent systems.
- Experience operating distributed systems with high availability and fault tolerance.
- Observability & Monitoring
- Hands‑on experience with monitoring and logging tools such as Prometheus, Grafana, ELK, or Loki.
- Understanding of incident management, service health metrics, and system reliability monitoring.
- DevOps & SRE Practices
- Solid understanding of SDLC, release processes, and operational reliability practices.
- Familiarity with CI/CD pipelines and Infrastructure‑as‑Code tools.
- Experience with GenAI, Agentic AI systems, or LLM orchestration frameworks.
- Exposure to LangChain, AutoGen, or RAG‑based systems.
- Experience with additional ML frameworks such as TensorFlow, JAX, or Ray.
- Knowledge of GPU/accelerator‑based systems and high‑performance networking (RDMA, InfiniBand, RoCE).
- Experience with advanced MLOps workflows or large‑scale AI platform operations.
- Bachelor’s or Master’s degree in engineering, Computer Science, AI/ML, or a related field.
- 12+ years of software, systems, or infrastructure engineering experience, preferably in production or datacenter environments.
Apart from working with great people, we offer the below:
- Salary including housing & transport allowance
- Stock (RSU's) and performance related bonus
- 16 weeks fully paid Maternity Leave
- 6 weeks fully paid Paternity Leave
- Employee stock purchase scheme
- Child Education Allowance
- Relocation and immigration support (if needed)
- Life and Medical Insurance
- Live+ Well Reimbursement for health and recreational membership fees
- Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 6+ years of Software Test Engineering or related work experience.
Master's degree in Engineering, Information Systems, Computer Science, or related field and 5+ years of Software Test Engineering or related work experience.
OR
PhD in Engineering, Information Systems, Computer Science, or related field and 4+ years of Software Test Engineering or related work experience.
- 2+ year of work experience with Software Test or System Test, developing and automating test plans, and/or tools (e.g., Source Code Control Systems, Continuous Integration Tools, and Bug Tracking Tools).
- References to a particular number of years experience are for indicative purposes only. Applications from candidates with equivalent experience will be considered, provided that the candidate can demonstrate an ability to fulfill the principal duties of the role and possesses the required competencies.
Qualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.
To all Staffing and Recruiting Agencies: Our Careers Site is only for individuals seeking a job at Qualcomm. Staffing and recruiting agencies and individuals being represented by an agency are not authorized to use this site or to submit profiles, applications or resumes, and any such submissions will be considered unsolicited. Qualcomm does not accept unsolicited resumes or applications from agencies. Please do not forward resumes to our jobs alias, Qualcomm employees or any other company location. Qualcomm is not responsible for any fees related to unsolicited resumes/applications.
If you would like more information about this role, please contact Qualcomm Careers.
المصدر: LinkedIn - أُضيفت للموقع في 13 أغسطس 2026
وظائف أخرى لدى Qualcomm