لوسيديا تعلن عن وظيفة مهندس موثوقية الموقع (SRE) في الرياض
تفاصيل الوظيفة
لوسيديا تبحث عن مهندس موثوقية الموقع (Site Reliability Engineer) للانضمام إلى فريقها في الرياض، المملكة العربية السعودية. ستكون مسؤولاً عن ضمان استقرار وأداء البنية التحتية السحابية لمنصة الذكاء الاصطناعي الخاصة بتجربة العملاء.
نبذة عن الوظيفة
في لوسيديا، تقوم المنصة بمعالجة كميات هائلة من بيانات العملاء في الوقت الفعلي. أي توقف أو تأخير أو عدم استقرار يؤثر بشكل مباشر على قدرة عملائنا على اتخاذ القرارات وخدمة مستخدميهم. هذا الدور موجود لضمان عدم حدوث ذلك. بصفتك مهندس موثوقية الموقع، ستكون في قلب استقرار المنصة، حيث تمتلك مسؤولية موثوقية البنية التحتية السحابية وضمان توسعها بسلاسة مع نمو الشركة. لن تتفاعل مع المشكلات فقط؛ بل ستتوقعها وتصمم أنظمة تمنعها وتبني أتمتة تزيلها بالكامل.
المهام والمسؤوليات
- تصميم وصيانة بنية تحتية عالية التوفر ومتسامحة مع الأخطاء وقابلة للتوسع.
- تحديد وإزالة نقاط الفشل الفردية بشكل استباقي قبل أن تتحول إلى حوادث.
- ضمان استقرار أنظمة الإنتاج تحت الأحمال المتزايدة.
- إدارة وتحسين أعباء العمل عبر AWS أو GCP أو Azure.
- استخدام البنية التحتية كرمز (IaC) باستخدام Terraform لتوحيد وتوسيع البنية التحتية.
- تحسين استخدام الموارد لتحقيق توازن بين الأداء والتكلفة.
- تشغيل وتوسيع مجموعات Kubernetes (مثل EKS, GKE) بثقة، واستكشاف الأخطاء وإصلاحها، وضمان نشر سلس.
- تنفيذ وتحسين أنظمة المراقبة باستخدام أدوات مثل Prometheus و Grafana و Datadog و ELK.
- تحديد تنبيهات ذات معنى (غير مزعجة) والاستجابة للحوادث وقيادة تحليل السبب الجذري.
- كتابة البرامج النصية (Python, Bash) وبناء أدوات لأتمتة العمل التشغيلي المتكرر.
- العمل مع فرق DevOps والهندسة لحل اختناقات الأداء والمساهمة في تحسينات CI/CD ونشر أفضل ممارسات الموثوقية.
- خلال أول 30 يوماً: بناء فهم قوي للبنية التحتية والأنظمة وسير العمل، والمساهمة في العمليات اليومية، وتحديد مجالات التحسين.
- بحلول 90 يوماً: إدارة مهام البنية التحتية بشكل مستقل، والمساهمة في تحسينات الموثوقية والتوسع، وامتلاك أجزاء من البنية التحتية وتحسينها.
الشروط والمتطلبات
- خبرة لا تقل عن 3 سنوات في مجال SRE أو DevOps أو هندسة البنية التحتية، مع فهم لكيفية حدوث الأعطال على نطاق واسع.
- إلمام ببيئات سحابية مثل AWS أو GCP أو Azure وفهم سلوك الأنظمة الموزعة.
- خبرة عملية مع Kubernetes في الإنتاج والقدرة على استكشاف الأخطاء وإصلاحها.
- عقلية تحليلية: لا تكتفي بإصلاح المشكلات بل تسأل عن سبب حدوثها وتضمن عدم تكرارها.
- خبرة في Terraform أو أدوات IaC مماثلة.
- خبرة في العمل مع Docker و Kubernetes.
- كتابة نصوص برمجية بلغة Python أو Bash أو ما شابه لأتمتة سير العمل.
- فهم قوي لخطوط أنابيب CI/CD (Jenkins, GitHub Actions, Bitbucket, إلخ).
- إلمام قوي بمفاهيم الشبكات وتوزيع الأحمال وتصميم التوافر العالي.
- خبرة في تنفيذ أدوات مراقبة مثل Prometheus و Grafana و Datadog و ELK، والقدرة على التمييز بين التنبيهات المفيدة والضجيج.
- روح المسؤولية: لا تنتظر حتى يقال لك أن شيئاً ما معطل.
- الهدوء تحت الضغط والمنهجية أثناء الحوادث.
- القدرة على تبسيط التعقيد بدلاً من زيادته.
- التواصل الواضح حتى عند شرح مشكلات تقنية عميقة.
- الاهتمام ببناء أنظمة تجعل المهندسين الآخرين أكثر فعالية.
- (إضافي - غير مطلوب) خبرة في RabbitMQ أو Redis في الإنتاج.
- (إضافي) معرفة بأدوات Ansible أو AWX.
- (إضافي) التعرض لبيئات متعددة السحابات أو هجينة.
- (إضافي) شهادات سحابية (AWS, GCP) أو شهادات Linux.
- (إضافي) خلفية من ITI (معهد تكنولوجيا المعلومات).
المهارات المطلوبة
- إتقان استخدام Terraform (أو أدوات IaC مماثلة) لإدارة البنية التحتية.
- العمل بثقة مع Docker و Kubernetes.
- كتابة نصوص برمجية بأتمتة سير العمل باستخدام Python أو Bash.
- فهم خطوط أنابيب CI/CD (Jenkins, GitHub Actions, Bitbucket).
- إلمام قوي بالشبكات وتوزيع الأحمال وتصميم التوافر العالي.
- تنفيذ أدوات المراقبة (Prometheus, Grafana, Datadog, ELK) والتركيز على الإشارات التي تقود العمل.
عرض النص الأصلي للإعلان
About Lucidya
Lucidya is an AI-native platform for customer experience (CX) intelligence that manages entire customer lifecycles autonomously, from initial engagement through retention and growth.
Unlike platforms that only surface insights and leave the action to you, Lucidya closes the loop with proprietary NLU technology built in-house and trained on millions of multilingual conversations. This enables marketing, support, CX, and research teams to deliver personalized experiences that drive measurable improvements in customer satisfaction, retention, and lifetime value.
As we continue scaling globally, the reliability, performance, and resilience of our infrastructure become mission-critical to everything we do.
Why this role matters
At Lucidya, our platform processes massive volumes of real-time customer data. Any downtime, latency, or instability directly impacts our customers’ ability to make decisions and serve their own users.
This role exists to make sure that doesn’t happen.
As a Site Reliability Engineer, you’ll sit at the heart of our platform’s stability, owning the reliability of our cloud infrastructure and ensuring it scales seamlessly as we grow. You won’t just react to issues; you’ll anticipate them, design systems that prevent them, and build automation that removes them entirely.
If you enjoy solving complex infrastructure challenges, eliminating inefficiencies, and building systems that “just work” - this is where you’ll thrive.
What You’ll Do
You’ll be responsible for outcomes, not just tasks. Here’s what success looks like in this role:
You’ll make reliability the default
- You’ll design and maintain infrastructure that is highly available, fault-tolerant, and scalable
- You’ll proactively identify and eliminate single points of failure before they become incidents
- You’ll ensure our production systems remain stable, even under increasing scale and load
You’ll own and optimize our cloud environments
- You’ll manage and continuously improve workloads across AWS, GCP, or Azure
- You’ll use Infrastructure as Code (Terraform) to standardize and scale infrastructure
- You’ll optimize resource usage to balance performance and cost
You’ll run and improve Kubernetes in production
- You’ll operate and scale Kubernetes clusters (EKS, GKE, etc.) with confidence
- You’ll troubleshoot issues quickly and ensure smooth deployments and upgrades
- You’ll ensure our containerized workloads perform reliably at scale
You’ll build strong observability and respond to incidents
- You’ll implement and refine monitoring systems using tools like Prometheus, Grafana, Datadog, or ELK
- You’ll define alerting that is meaningful, not noisy
- You’ll respond to incidents, lead root cause analysis, and ensure we learn from every failure
You’ll automate everything that shouldn’t be manual
- You’ll write scripts and build tooling to eliminate repetitive operational work
- You’ll continuously improve infrastructure efficiency through automation
- You’ll promote a culture where manual work is a temporary state, not the norm
You’ll collaborate to improve the entire system
- You’ll work closely with DevOps and engineering teams to solve performance bottlenecks
- You’ll contribute to CI/CD improvements and deployment reliability
- You’ll help shape reliability best practices across the organization
What success looks like (First 90 Days)
First 30 days:
- You’ve built a strong understanding of our infrastructure, systems, and workflows
- You’re contributing to day-to-day operations with support from the team
- You’ve started identifying areas for improvement in automation and reliability
By 90 days:
- You’re independently managing infrastructure tasks and troubleshooting issues
- You’re actively contributing to reliability and scalability improvements
- You’ve taken ownership of parts of our infrastructure and are improving them
Requirements
Who You Are
This is what will make you successful in this role:
- You’ve spent ~3 years working in SRE, DevOps, or infrastructure engineering, and you’ve seen what breaks at scale
- You’re comfortable working in cloud environments like AWS, GCP, or Azure-and you understand how distributed systems behave
- You’ve worked hands-on with Kubernetes in production and know how to troubleshoot it when things go wrong
- You don’t just fix issues - you ask why they happened and make sure they don’t happen again
Technically, you likely:
- Use Terraform (or similar IaC tools) to manage infrastructure
- Work confidently with Docker and Kubernetes
- Write scripts in Python, Bash, or similar to automate workflows
- Understand CI/CD pipelines (Jenkins, GitHub Actions, Bitbucket, etc.)
- Have a solid grasp of networking, load balancing, and high-availability design
When it comes to monitoring:
- You’ve implemented tools like Prometheus, Grafana, Datadog, or ELK
- You know the difference between useful alerts and noise
- You focus on signals that actually drive action
What sets you apart:
- You take ownership - you don’t wait to be told something is broken
- You’re calm under pressure and methodical during incidents
- You simplify complexity instead of adding to it
- You communicate clearly, even when explaining deeply technical issues
- You care about building systems that make other engineers more effective
Nice to Have (but not required)
- Experience with RabbitMQ or Redis in production
- Familiarity with Ansible or AWX
- Exposure to multi-cloud or hybrid environments
- Cloud certifications (AWS, GCP) or Linux certifications
- Background from ITI (Information Technology Institute)
What the hiring process will look like
- Screening Interview - Talent Acquisition
- Technical Interview - SRE Lead
- Technical Task
- Final Interview - SRE Lead & Cloud DevOps Director
وظائف أخرى لدى لوسيديا