SmartChoice International GCC تعلن عن وظيفة مهندس DevOps أول في الرياض
تفاصيل الوظيفة
بالتعاون مع إحدى الشركات التقنية المتنامية، يعلن مكتب SmartChoice International GCC عن توفر وظيفة Senior DevOps Engineer في الرياض، السعودية. ستكون مسؤولاً عن بناء وتشغيل البنية التحتية لمنصة تقنية سريعة التوسع، مع التركيز على السحابة، الأتمتة، CI/CD، الحاويات، الأمان، المراقبة، وموثوقية الإنتاج.
المهام والمسؤوليات
- تصميم وصيانة بنية تحتية قابلة للتكرار باستخدام البنية التحتية كرمز (IaC).
- إدارة خطوط CI/CD شاملة لمراحل البناء والاختبار والأمان والنشر الإنتاجي.
- إدارة أحمال العمل المحوسبة عبر Kubernetes وبيئات الحوسبة السحابية.
- بناء استراتيجيات نشر آمنة مع آليات طرح وتراجع واستعادة مناسبة.
- إدارة الشبكات السحابية، بوابات API، موازنة التحميل، DNS، الشهادات، واتصال الخدمات.
- تنفيذ إدارة الهوية، إدارة الأسرار، والوصول بأقل صلاحية عبر البيئات.
- تشغيل قواعد بيانات إنتاجية بما يشمل التوفر العالي، التكرار، النسخ الاحتياطي والاستعادة.
- إنشاء أنظمة مراقبة وتنبيه ومقاييس تشغيلية للبنية التحتية والتطبيقات.
- قيادة الاستجابة للحوادث وتحسين الموثوقية والأمان والمرونة التشغيلية باستمرار.
- تحسين استخدام البنية التحتية وتكاليفها، بما في ذلك أحمال العمل عالية الحوسبة.
- المساعدة في إنشاء قدرات منصة قابلة لإعادة الاستخدام لتمكين فرق التطوير من نشر وتشغيل الخدمات بشكل مستقل.
الشروط والمتطلبات
- 8+ سنوات من الخبرة في DevOps أو SRE أو البنية التحتية أو هندسة المنصات.
- خبرة عملية قوية مع AWS تشمل الحوسبة، الشبكات، IAM، والتخزين.
- خبرة قوية في البنية التحتية كرمز باستخدام Terraform أو OpenTofu، بما يشمل الوحدات القابلة لإعادة الاستخدام وإدارة الحالة (state).
- خبرة مع Ansible أو أدوات إدارة التهيئة المماثلة.
- خبرة قوية في CI/CD وخاصة GitHub Actions أو ما يعادلها.
- خبرة قوية مع Docker و Kubernetes و Helm، ويفضل خبرة مع EKS.
- خبرة مع GitOps بما يشمل ArgoCD أو Flux.
- خبرة مع بوابات API، المصادقة، التوجيه، وتحديد المعدل (rate limiting).
- خبرة إنتاجية مع PostgreSQL أو RDS/Aurora أو Redis أو تقنيات مماثلة.
- خبرة مع Vault أو AWS Secrets Manager أو Parameter Store.
- خبرة قوية في المراقبة مع Prometheus أو Grafana أو CloudWatch أو OpenTelemetry أو ما يعادلها.
- ملمّ بالعمل في بيئات Linux واستخدام Bash و Python.
- فهم قوي للأمان، الموثوقية، إدارة الحوادث، وعمليات الإنتاج.
المهارات المطلوبة
يُفضّل خبرة متقدمة في تشغيل أعباء عمل الذكاء الاصطناعي أو التعلم الآلي أو غيرها من الأحمال عالية الحوسبة في الإنتاج، مثل: البنية التحتية لوحدات معالجة الرسوميات (GPU) وجدولة أعباء العمل، منصات تقديم النماذج (model-serving) أو الاستدلال (inference)، التحجيم التلقائي وإدارة السعة للأحمال المكثفة حاسوبياً، إدارة أداء واستخدام وتكاليف البيئات عالية الحوسبة. التقنيات ذات الصلة قد تشمل: vLLM, Triton, TGI, KServe, Ray Serve, Amazon Bedrock, أو SageMaker. البيئة التقنية العامة: AWS, Terraform/OpenTofu, Ansible, GitHub Actions, Docker, Kubernetes/EKS, Helm, ArgoCD/Flux, API Gateways, PostgreSQL, RDS/Aurora, Redis, Vault/Secrets Manager, Prometheus, Grafana, CloudWatch, OpenTelemetry, Linux, Bash, Python.
عرض النص الأصلي للإعلان
Senior DevOps Engineer
Full-time | Riyadh
We are partnering with a growing technology organisation to appoint a Senior DevOps Engineer to build and operate the infrastructure behind a rapidly scaling technology platform.
This is a hands-on role spanning cloud infrastructure, automation, CI/CD, container platforms, security, observability and production reliability. You will work closely with engineering teams to create scalable, secure and resilient environments, while establishing the operational standards the wider technology function can build on.
Key Responsibilities
- Design and maintain scalable, reproducible infrastructure using infrastructure as code.
- Own CI/CD pipelines covering build, testing, security and production deployment.
- Manage containerised workloads across Kubernetes and cloud-based compute environments.
- Build safe deployment strategies with appropriate rollout, rollback and recovery mechanisms.
- Manage cloud networking, API gateways, load balancing, DNS, certificates and service connectivity.
- Implement identity, secrets management and least-privilege access across environments.
- Operate production databases, including high availability, replication, backup and recovery.
- Establish monitoring, alerting and operational metrics across infrastructure and applications.
- Lead incident response and continuously improve reliability, security and operational resilience.
- Optimise infrastructure utilisation and costs, including high-compute workloads.
- Help create reusable platform capabilities that enable engineering teams to deploy and operate services independently.
What We're Looking For
- 8+ years of experience in DevOps, SRE, infrastructure or platform engineering.
- Strong hands-on experience with AWS, including compute, networking, IAM and storage.
- Strong infrastructure-as-code experience with Terraform or OpenTofu, including reusable modules and state management.
- Experience with Ansible or comparable configuration management.
- Strong CI/CD experience, particularly GitHub Actions or equivalent.
- Strong experience with Docker, Kubernetes and Helm, with EKS highly desirable.
- Experience with GitOps, including ArgoCD or Flux.
- Experience with API gateways, authentication, routing and rate limiting.
- Production experience with PostgreSQL, RDS/Aurora and Redis or equivalent technologies.
- Experience with Vault, AWS Secrets Manager or Parameter Store.
- Strong observability experience with Prometheus, Grafana, CloudWatch and OpenTelemetry or equivalents.
- Comfortable working in Linux environments with Bash and Python.
- Strong understanding of security, reliability, incident management and production operations.
Advanced Infrastructure
Experience operating AI, machine learning or other high-compute workloads in production is highly desirable.
Relevant experience may include:
- GPU infrastructure and workload scheduling.
- Model-serving or inference platforms.
- Autoscaling and capacity management for compute-intensive workloads.
- Managing performance, utilisation and cost of high-compute environments.
- Technologies such as vLLM, Triton, TGI, KServe, Ray Serve, Amazon Bedrock or SageMaker.
Technical Environment
AWS | Terraform/OpenTofu | Ansible | GitHub Actions | Docker | Kubernetes/EKS | Helm | ArgoCD/Flux | API Gateways | PostgreSQL | RDS/Aurora | Redis | Vault/Secrets Manager | Prometheus | Grafana | CloudWatch | OpenTelemetry | Linux | Bash | Python
We're looking for an engineer with strong technical depth and operational judgement who can work calmly through complex production problems, understand the impact of infrastructure changes, and build systems with reliability, security and cost in mind.
رقم الإعلان لدى المصدر: 4460728084