About TestCrew and the Role
TestCrew | Quality Engineering & Software Testing is seeking a Site Reliability Engineer (SRE) for an upcoming enterprise engagement in Al-Ahsa, located in the Eastern region. This full-time position is open to candidates with 2-5 years of relevant experience, focusing on ensuring the stability and performance of critical IT systems.
Role Overview
The Site Reliability Engineer will undertake a hands-on operational role, responsible for the availability, reliability, performance, and security of critical IT systems. This involves proactive monitoring, automation, and disciplined incident response. The successful candidate will collaborate closely with development, operations, and security teams to maintain highly available services, streamline operations through automation, and drive continuous service improvement.
Key Responsibilities
- Monitor and manage enterprise infrastructure, applications, and services to ensure high availability, stability, and optimal performance.
- Operate and maintain monitoring and observability platforms to detect incidents, analyze trends, and respond proactively to system issues.
- Design and implement monitoring strategies, dashboards, alerts, and performance metrics to improve operational visibility.
- Perform root cause analysis (RCA) for incidents and implement preventive measures to reduce recurring issues.
- Automate operational tasks, deployments, and maintenance activities using scripting and infrastructure automation tools.
- Manage CI/CD pipelines and support release management processes to enable reliable and efficient software deployments.
- Optimize system performance while ensuring compliance with service level agreements (SLAs), operational standards, and best practices.
- Support business continuity, disaster recovery, and operational resilience initiatives.
- Implement security best practices, including system hardening, patch management, and secure operational procedures.
- Collaborate with development, infrastructure, security, and IT operations teams to resolve complex technical issues.
- Maintain operational documentation, runbooks, knowledge articles, and incident reports.
- Participate in major incident management, post-incident reviews, and continuous improvement initiatives.
Candidate Profile and Requirements
- 2-5 years of experience in a Site Reliability Engineering or a similar operational role.
- Strong ownership mindset with a proactive, reliability-first approach.
- Excellent analytical and problem-solving skills.
- Ability to remain calm and methodical during critical incidents.
- Strong communication and collaboration skills.
- Ability to work effectively in a fast-paced enterprise environment.
Work Environment and Conditions
This role requires candidates to work onsite at the client location in Al-Ahsa. Participation in on-call support and shift rotations is also a requirement for this position.
Application Information
This full-time position offers an opportunity to contribute to critical system reliability within an enterprise environment. Salary details will be discussed during the interview process.
عن TestCrew والدور الوظيفي
TestCrew | الهندسة الجودة واختبار البرمجيات يسعى لتعيين مهندس موثوقية المواقع (SRE) لمشاركة مؤسسية قادمة في الأحساء، الواقعة في المنطقة الشرقية. هذا المنصب بدوام كامل مفتوح للمرشحين الذين لديهم خبرة ذي صلة من 2-5 سنوات، مع التركيز على ضمان استقرار وأداء أنظمة تقنية المعلومات الحيوية.
نظرة عامة على الدور
سيشغل مهندس موثوقية المواقع دوراً عملياً تشغيلياً، مسؤول عن التوفر، والموثوقية، والأداء، وأمان الأنظمة الحيوية لتقنية المعلومات. يشمل ذلك المراقبة الاستباقية، الأتمتة، والاستجابة للحدث بشكل منضبط. سيتعاون المرشح الناجح بشكل وثيق مع فرق التطوير، والعمليات، والأمن للحفاظ على خدمات عالية التوفر، وتبسيط العمليات من خلال الأتمتة، ودفع التحسين المستمر للخدمات.
المسؤوليات الرئيسية
- مراقبة وإدارة بنية المؤسسة والتطبيقات والخدمات لضمان التوفر العالي، والاستقرار، والأداء الأمثل.
- تشغيل وصيانة منصات المراقبة والملاحظية لاكتشاف الحوادث، وتحليل الاتجاهات، والرد بشكل استباقي على مشكلات النظام.
- تصميم وتنفيذ استراتيجيات المراقبة ولوحات البيانات والتنبيهات ومقاييس الأداء لتحسين الرؤية التشغيلية.
- إجراء تحليل السبب الجذري (RCA) للحوادث وتنفيذ إجراءات وقائية لتقليل تكرار المشكلات.
- أتمتة المهام التشغيلية، ونشر التحديثات، وأنشطة الصيانة باستخدام أدوات البرمجة وأتمتة البنية التحتية.
- إدارة خطوط أنابيب CI/CD ودعم عمليات إدارة الإصدار لتمكين نشر البرمجيات بشكل موثوق وفعّال.
- تحسين أداء النظام مع ضمان الالتزام باتفاقيات مستوى الخدمة (SLAs)، والمعايير التشغيلية، وأفضل الممارسات.
- دعم استمرارية الأعمال، والاسترداد من الكوارث، ومبادرات المرونة التشغيلية.
- تنفيذ أفضل ممارسات الأمن، بما في ذلك تقوية الأنظمة، إدارة التصحيحات، وإجراءات تشغيل آمنة.
- التعاون مع فرق التطوير والبنية التحتية والأمن وعمليات تقنية المعلومات لحل قضايا تقنية معقدة.
- الحفاظ على الوثائق التشغيلية ودليل التشغيل ومقالات المعرفة وتقارير الحوادث.
- المشاركة في إدارة الحوادث الكبرى، ومراجعات ما بعد الحادث، ومبادرات التحسين المستمر.
ملف المرشح والمتطلبات
- 2-5 سنوات من الخبرة في هندسة موثوقية المواقع أو دور عملي مشابه.
- عقلية امتلاك قوية مع نهج استباقي يعتمد الأولوية على الاعتمادية.
- مهارات تحليلية وحل مشاكل ممتازة.
- القدرة على الحفاظ على الهدوء والمنهجية خلال الحوادث الحرجة.
- مهارات تواصل وتعاون قوية.
- القدرة على العمل بفاعلية في بيئة مؤسسية سريعة الوتيرة.
بيئة العمل والظروف
يتطلب هذا الدور وجود المرشحين للعمل في موقع العميل في الأحساء. كما أن المشاركة في الدعم المتواصل وتناوب النوبات مطلوبان لهذا المنصب.
معلومات التقديم
يوفر هذا المنصب بدوام كامل فرصة للمساهمة في موثوقية النظام الحيوي داخل بيئة المؤسسة. ستتم مناقشة تفاصيل الراتب أثناء عملية المقابلة.