وظيفة Site Reliability Engineer شاغرة لدى FNRCO في الدمام السعودية
تفاصيل الوظيفة
نحن في FNRCO نبحث عن مهندس Site Reliability Engineer (SRE) للعمل في الشرقية، الدمام، السعودية. يتطلب الدور خبرة من 5 إلى 15 سنة في DevSecOps وSRE لتصميم وتشغيل وتأمين بيئات سحابية منظمة على AWS وGCP، مع التركيز على Kubernetes وTerraform وCI/CD وHashiCorp Vault وDatadog وPagerDuty.
المهام والمسؤوليات
- نشر وإدارة وصيانة استراتيجية السحابة السيادية (Sovereign Cloud)، وضمان الامتثال للمتطلبات التنظيمية ومتطلبات إقامة البيانات على AWS/GCP.
- إدارة وتشغيل مجموعات Kubernetes بما في ذلك الترقيات والتوسع وتحسين أعباء العمل.
- الاستجابة للحوادث الأمنية وتخفيف المخاطر: التحقيق في الحوادث الأمنية السحابية وتحليلها ومعالجتها، وتحديد الثغرات والتخفيف منها في بيئات AWS.
- دعم وتعزيز استراتيجية AWS السحابية من خلال تضمين أفضل الممارسات الأمنية وتحسين الوضع الأمني بشكل مستمر.
- تنفيذ وضبط أدوات الأمان، وأتمتة الاستجابات القائمة على السياسات، والدعوة لممارسات DevSecOps لضمان عمليات سحابية آمنة بالتصميم.
- تنفيذ وإدارة البنية التحتية كرمز (Terraform) لتوفير وتعديل وتأمين الموارد السحابية.
- صيانة وتحسين خطوط أنابيب CI/CD باستخدام Git/GitLab، وضمان عمليات نشر آمنة ومؤتمتة.
- إدارة الأسرار والوصول الآمن باستخدام HashiCorp Vault، بما في ذلك دورة حياة الرموز وسياسات الوصول وتدوير الأسرار.
- استكشاف المشكلات المعقدة في البنية التحتية والشبكات والحاويات والأداء عبر الأنظمة الموزعة وإصلاحها.
- مراقبة صحة النظام وأدائه باستخدام Datadog، وتحديد وإدارة SLIs وSLOs، ودفع التحسينات المستمرة للموثوقية وفقًا لمبادئ SRE.
- إدارة التنبيهات على مدار الساعة طوال أيام الأسبوع والاستجابة للحوادث عبر PagerDuty، وإجراء تحليل السبب الجذري (RCA)، والمشاركة بنشاط في عمليات إدارة الحوادث والمشكلات والتغيير.
الشروط والمتطلبات
- خبرة عملية من 5 إلى 15 سنة في أدوار DevSecOps أو SRE.
- خبرة عملية مؤكدة في إدارة مجموعات Kubernetes.
- خبرة في نشر Terraform (IaC) واستخدام Git/GitLab لـ CI/CD والتحكم في الإصدارات.
- خبرة في العمل مع HashiCorp Vault وTerraform وDatadog وPagerDuty وConfluence وغيرها.
- فهم قوي لمبادئ أمن السحابة (IAM، التشفير، أمن الشبكات، أمن الحاويات، إدارة الثغرات).
- خبرة في عمليات إدارة الحوادث وإدارة التغيير وتحليل السبب الجذري.
- خبرة في مبادئ SRE (SLIs، SLOs، ميزانيات الأخطاء، مقاييس الموثوقية) مع خبرة عملية في البرمجة النصية بلغة Python أو Bash.
- إلمام قوي بأساسيات الشبكات (TCP/IP، DNS، موازنة التحميل، جدران الحماية، VPNs، نقاط النهاية الخاصة) مع خبرة أو تعرض لبيئات منظمة وملتزمة بالامتثال.
- القدرة على قيادة فرق أثناء انقطاعات P1/P2، وتنسيق الفرق متعددة الوظائف ودفع الحلول في الوقت المناسب مع تحليل سبب جذري واضح.
- الحفاظ على عقلية تركز على العملاء والأعمال مع تحديد أولويات المهام.
- الشهادات السحابية (مثل AWS أو GCP) تعتبر ميزة إضافية.
المهارات المطلوبة
- إدارة Kubernetes وتشغيله (ترقيات، توسع، تحسين أعباء العمل).
- البنية التحتية كرمز باستخدام Terraform.
- إدارة CI/CD باستخدام Git/GitLab.
- إدارة الأسرار باستخدام HashiCorp Vault.
- المراقبة والتنبيهات باستخدام Datadog وPagerDuty.
- البرمجة النصية بلغة Python أو Bash للأتمتة.
- أتمتة البنية التحتية باستخدام Ansible.
- فهم متقدم لأمن السحابة (IAM، التشفير، أمن الشبكات، أمن الحاويات).
- أساسيات الشبكات (TCP/IP، DNS، موازنة التحميل، جدران الحماية، VPNs).
- مبادئ SRE (SLIs، SLOs، ميزانيات الأخطاء).
- إدارة الحوادث وتحليل السبب الجذري (RCA).
- العمل في بيئات سحابية منظمة وملتزمة بالامتثال (AWS/GCP).
عرض النص الأصلي للإعلان
Job Overview:
We are seeking a skilled DevSecOps / Sovereign Cloud Engineer with 5-15 years of experience to design, operate, and secure regulated cloud environments across AWS and GCP. The role combines cloud operations, security engineering, and SRE practices, including Kubernetes management, Infrastructure as Code (Terraform), CI/CD automation (Git/GitLab), and secrets management (HashiCorp Vault). The ideal candidate will drive security incident response, platform reliability, observability (Datadog), and 24x7 incident management (PagerDuty), while embedding security-by-design principles, automation-first practices, and continuous improvement into sovereign cloud platforms. Strong expertise in cloud security, networking fundamentals, compliance-driven environments, and cross-functional incident leadership is essential.
Key Responsibilities:
- Deploy, manage and maintain the organization’s Sovereign Cloud strategy, ensuring compliance with regulatory and data residency requirements on AWS/GCP public cloud.
- Managing and operating Kubernetes clusters including upgrades, scaling, and workload optimization.
- Security Incident Response & Risk Mitigation - Investigate, analyze, and remediate cloud security incidents, proactively identifying and mitigating vulnerabilities within AWS environments.
- Secure Cloud Strategy & Continuous Improvement - Support and enhance the organization’s AWS cloud strategy by embedding security best practices and continuously improving the cloud security posture.
- Security Automation & DevSecOps Enablement - Implement and tune security tools, automate policy-driven responses, and advocate DevSecOps practices to ensure secure-by-design cloud operations.
- Implementing and managing Infrastructure as Code (Terraform) to provision, modify, and secure cloud resources.
- Maintaining and optimizing CI/CD pipelines using Git/GitLab, ensuring secure and automated deployments.
- Managing secrets and secure access using HashiCorp Vault, including token lifecycle, access policies, and secrets rotation.
- Troubleshooting complex infrastructure, networking, container, and performance issues across distributed systems.
- Observability & Reliability Engineering - Monitor system health and performance using Datadog, define and manage SLIs/SLOs, and drive continuous reliability improvements aligned with SRE principles.
- Incident Management & Operational Governance - Manage 24x7 alerting and incident response through PagerDuty, perform root cause analysis (RCA), and actively contribute to incident, problem, and change management processes.
- Cloud Security & Performance Optimization - Conduct proactive system hardening, vulnerability remediation, performance tuning, and capacity planning across cloud environments.
- Automation & Continuous Improvement - Develop automation using Python/Bash/Terraform/Ansible to reduce manual effort, improve operational efficiency, and strengthen platform resilience.
Knowledge, Skills and Experience:
To succeed in this role, you must have:
- 5-15 years of hands-on experience in DevSecOps, SRE roles.
- Proven hands-on expertise managing Kubernetes clusters
- Experience with Terraform (IaC) Deployment, Git/Gitlab for CI/CD and version control for best practices
- Experience working with HashiCorp Vault, Terraform, Datadog, Pagerduty, Confluence etc
- Strong understanding of Cloud Security principles (IAM, encryption, network security, container security, vulnerability management).
- Experience in incident management, change management, and root cause analysis processes.
- SRE & Automation Expertise - Strong understanding of SRE principles (SLIs, SLOs, error budgets, reliability metrics) with hands-on scripting experience in Python and/or Bash to drive automation-first practices.
- Networking & Compliance Knowledge - Solid grasp of networking fundamentals (TCP/IP, DNS, load balancing, firewalls, VPNs, private endpoints) with experience or exposure to regulated and compliance-driven environments.
- Leads incident bridges during P1/P2 outages, coordinating cross-functional teams and driving timely resolution with clear RCA.
- Maintains a strong customer and business-focused mindset while prioritizing tasks.
- Cloud certifications in a plus.
Also, You can forward your CV through below link for more upcoming Job vacancies:
https://cv-fnrco.com