About the Role at TawanTech
TawanTech is seeking a Senior Microservices Support Engineer to join our team in Riyadh, Riyadh Province. This full-time position involves leading the production support and ensuring the operational excellence of mission-critical microservices applications. The role is integral to a 24x7 support organization, requiring participation in rotating shifts and on-call coverage to maintain high availability and stability for enterprise systems.
Role Context and Objectives
The primary objective of this role is to ensure rapid incident resolution and proactive system monitoring for complex enterprise environments. The ideal candidate will apply deep expertise in Java-based microservices, OpenShift, advanced troubleshooting, and DevOps practices to drive continuous service improvement and uphold stringent service level agreements.
Key Responsibilities
- Lead Level 2/Level 3 production support for enterprise microservices applications.
- Act as the technical lead during production incidents, coordinating resolution with development, infrastructure, and DevOps teams.
- Perform advanced troubleshooting and root cause analysis for application, middleware, infrastructure, and performance-related issues.
- Proactively monitor production systems to identify risks, performance degradation, and potential outages.
- Support application deployments, release management, rollbacks, and environment readiness within OpenShift.
- Ensure compliance with SLAs by effectively managing incidents, service requests, and problem tickets.
- Analyze application logs, JVM behavior, container health, and infrastructure metrics to resolve complex issues.
- Drive continuous service improvements by identifying recurring issues and implementing permanent solutions.
- Develop and maintain operational runbooks, knowledge base articles, and support documentation.
- Mentor and provide technical guidance to junior support engineers.
- Collaborate with development teams to improve application reliability, observability, and supportability.
- Participate in release planning, production readiness reviews, and post-implementation validation.
- Support disaster recovery activities, failover testing, and business continuity initiatives.
Required Qualifications and Experience
- Minimum of 6–10+ years of experience in Application Support, Production Support, or Site Reliability Engineering (SRE).
- At least 4+ years of experience supporting Java-based microservices in enterprise production environments.
- Hands-on experience with OpenShift or Kubernetes platforms.
- Proven experience supporting highly available, mission-critical applications.
- Strong understanding of distributed systems, caching strategies, and messaging platforms.
- Experience working within Agile and DevOps environments.
Preferred Technical Skills
- Familiarity with cloud platforms such as AWS, Azure, or GCP is considered an advantage.
Work Schedule and Commitment
This is a full-time role that requires participation in a 24x7 support organization. The successful candidate will be expected to work rotating shifts and participate in an on-call support rotation for critical production systems.
عن الدور في TawanTech
توان تكنولوجي تسعى إلى مهندس دعم الخدمات المصغرة الأقدم للانضمام إلى فريقنا في الرياض، منطقة الرياض. يتضمن هذا المنصب بدوام كامل قيادة دعم الإنتاج وضمان التميز التشغيلي لتطبيقات الخدمات المصغرة الحيوية. الدور جزء من منظومة دعم على مدار 24x7، ويتطلب المشاركة في نوبات متغيرة والتغطية عند الطلب للحفاظ على التوفر العالي والاستقرار للأنظمة المؤسسية.
سياق الدور والأهداف
الهدف الأساسي من هذا الدور هو ضمان حل الحوادث بسرعة والمراقبة الاستباقية للنظم لبيئات المؤسسات المعقدة. سيطبق المرشح المثالي خبرة عميقة في الخدمات المصغرة المعتمدة على Java، OpenShift، استكشاف الأخطاء المتقدم، وممارسات DevOps لدفع تحسين الخدمة المستمر والحفاظ على اتفاقيات مستوى الخدمة الصارمة.
المسؤوليات الأساسية
- قيادة دعم الإنتاج من المستوى 2/المستوى 3 لتطبيقات الخدمات المصغرة المؤسسية.
- أن تكون القائد التقني أثناء حوادث الإنتاج، وتنسيق الحلول مع فرق التطوير والبنية التحتية وDevOps.
- إجراء استكشاف أخطاء متقدم وتحليل السبب الجذري لقضايا التطبيق والوسيط والبنية التحتية والأداء.
- مراقبة النظم الإنتاجية بشكل استباقي لتحديد المخاطر وتدهور الأداء واحتمالات الانقطاع.
- دعم نشر التطبيقات وإدارة الإصدارات والتراجع وتهيئة البيئات ضمن OpenShift.
- ضمان الالتزام باتفاقيات مستوى الخدمة من خلال إدارة الحوادث وطلبات الخدمة وتذاكر المشكلة بفعالية.
- تحليل سجلات التطبيق، سلوك JVM، صحة الحاويات، وقياسات البنية التحتية لحل القضايا المعقدة.
- قيادة تحسينات الخدمة المستمرة من خلال تحديد القضايا المتكررة وتنفيذ حلول دائمة.
- تطوير وصيانة كتيبات تشغيلية، مقالات قاعدة المعرفة، ومستندات الدعم.
- توجيه وتقديم الإرشاد التقني لمهندسي الدعم المبتدئين.
- التعاون مع فرق التطوير لتحسين موثوقية التطبيق والمراقبة والقدرة على الدعم.
- المشاركة في تخطيط الإصدارات ومراجعات جاهزية الإنتاج والتحقق بعد التنفيذ.
- دعم أنشطة التعافي من الكوارث، اختبارات الفشل، ومبادرات استمرارية الأعمال.
المؤهلات والخبرة المطلوبة
- على الأقل 6–10+ سنوات من الخبرة في دعم التطبيقات، دعم الإنتاج، أو هندسة موثوقية المواقع (SRE).
- على الأقل 4+ سنوات من الخبرة في دعم الخدمات المصغرة المعتمدة على Java في بيئات إنتاج المؤسسة.
- خبرة عملية مع منصات OpenShift أو Kubernetes.
- خبرة مثبتة في دعم تطبيقات عالية التAvailability والحيوية.
- فهم قوي للأنظمة الموزعة، واستراتيجيات التخزين المؤقت، ومنصات الرسائل.
- خبرة في العمل ضمن بيئات Agile وDevOps.
المهارات التقنية المفضلة
- المألوفية مع منصات السحابة مثل AWS أو Azure أو GCP تعتبر ميزة.
جدول العمل والالتزام
هذا دور بدوام كامل يتطلب المشاركة في منظمة دعم تعمل على مدار 24x7. من المتوقع أن يعمل المرشح الناجح في نوبات متغيرة ويشارك في تدوير الدعم عند الطلب للنظم الإنتاجية الحرجة.