Mindrift is looking for highly skilled Web Scraping specialists to join the Tendem project ( https://tendem.ai/ ) and drive specialized data scraping workflows for real-world use cases. Mindrift is looking for highly skilled Python Data Scraping Engineers to join the Tendem project and drive specialized data scraping workflows for real-world applications. In this role, you'll apply your expertise in web scraping, data extraction, and data processing to deliver accurate, reliable, and high-quality results. This part-time remote opportunity is ideal for technical professionals with hands-on experience in web scraping, data extraction, and processing.
This is a freelance role for a Tendem project. As a Python Data Scraping Engineer, you'll handle data scraping tasks requiring technical precision for web extraction and processing, utilizing tools such as Apify, OpenRouter, and other technologies, alongside your own technical expertise and approaches.
Key Responsibilities
- Own end-to-end data extraction workflows across complex websites, ensuring complete coverage, accuracy, and reliable delivery of structured datasets.
- Leverage available tools and custom workflows to accelerate data collection, validation, and task execution while meeting defined requirements.
- Ensure reliable extraction from dynamic and interactive web sources, adapting approaches as needed to handle JavaScript-rendered content and changing site behavior.
- Enforce data quality standards through validation checks, cross-source consistency controls, adherence to formatting specifications, and systematic verification prior to delivery.
- Scale scraping operations for large datasets using efficient batching or parallelization, monitor failures, and maintain stability against minor site structure changes.
Technical Skills (Essential)
- Strong experience in Python web scraping (BeautifulSoup, Selenium or similar), including dynamic content (JS, AJAX, infinite scroll) and APIs via proxies
- Proven ability to extract data from complex structures (hierarchies, archived pages, inconsistent HTML)
- Solid background in data cleaning, normalization, and validation, delivering structured datasets (CSV, JSON, Google Sheets)
Additional requirements
- Hands-on experience with LLMs and AI frameworks to enhance automation and problem-solving
- Strong attention to detail and commitment to data accuracy
- Self-directed work ethic with ability to troubleshoot independently
- A link to GitHub is a plus
- English proficiency: Upper-intermediate (B2) or above (required)
For this project, tasks are estimated to require around 10 20 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active.
On this project, contributors can earn up to $37 per hour equivalent , depending on their level and pace of contribution. Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.
Desired Candidate Profile
Educational qualifications
At least 3+ years of relevant experience in data engineering, web scraping, automation, or software development (required).
Bachelor s or Master s Degree in Engineering, Applied Mathematics, Computer Science, or related technical fields is a plus.
Academic and/or Professional Experience
Candidates should have a strong technical foundation and practical experience with scripting, automation, and data extraction workflows. We are looking for specialists who can solve non-trivial problems, work confidently with modern web technologies and data processing tools, and systematically collect, structure, and validate data from diverse sources. A methodical, detail-oriented approach and the ability to work independently are essential.
تبحث Mindrift عن متخصصين عاليي المهارة في كشط الويب للانضمام إلى مشروع Tendem ( https://tendem.ai/ ) ودفع سير عمل كشط بيانات مخصصة لحالات استخدام واقعية. تبحث Mindrift عن مهندسي كشط بيانات بايثون عاليي المهارة للانضمام إلى مشروع Tendem ودفع سير عمل كشط بيانات مخصصة لتطبيقات واقعية. في هذا الدور، ستطبق خبرتك في كشط الويب، واستخراج البيانات، ومعالجتها لتقديم نتائج دقيقة وموثوقة وعالية الجودة. هذه الفرصة عن بُعد بدوام جزئي مثالية للمهنيين التقنيين ذوي الخبرة العملية في كشط الويب، واستخراج البيانات، والمعالجة.
هذه وظيفة حُرّة لمشروع Tendem. كمهندس كشط بيانات بلغة Python، ستتعامل مع مهام كشط البيانات التي تتطلب دقة تقنية في الاستخراج والمعالجة على الويب، مستخدمًا أدوات مثل Apify وOpenRouter وتكنولوجيات أخرى، إلى جانب خبرتك ومناهجك التقنية الخاصة.
المسؤوليات الأساسية
- امتلاك تدفقات استخراج البيانات من البداية وحتى النهاية عبر مواقع معقدة، مع ضمان التغطية الكاملة والدقة والتسليم الموثوق للمجموعات البيانات المهيكلة.
- استغلال الأدوات المتاحة وسير العمل المخصص لتسريع جمع البيانات، والتحقق منها، وتنفيذ المهام مع الالتزام بالمتطلبات المحددة.
- ضمان استخراج موثوق من مصادر ويب ديناميكية وتفاعلية، مع تكييف الأساليب حسب الحاجة للتعامل مع المحتوى المعروض جافا سكريبت وتغير سلوك الموقع.
- فرض معايير جودة البيانات من خلال فحوصات التحقق، ضوابط الاتساق عبر المصادر، الالتزام بمواصفات التنسيق، والتحقق المنهجي قبل التسليم.
- زيادة نطاق عمليات الكشط لبيانات كبيرة باستخدام تجميع فعال أو تعدد المعالجات، رصد الإخفاقات، والحفاظ على الاستقرار أمام تغيّر بسيط في بنية الموقع.
المهارات التقنية (أساسية)
- خبرة قوية في كشط الويب باستخدام بايثون (BeautifulSoup، Selenium أو ما يماثلها)، بما في ذلك المحتوى الديناميكي (JS، AJAX، التمرير اللانهائي) وواجهات برمجة التطبيقات عبر البروكسيات
- قدرة مثبتة على استخراج البيانات من تراكيب معقدة (الهياكل، الصفحات المؤرشفة، HTML غير المتسق)
- خلفية راسخة في تنظيف البيانات، التطبيع، والتحقق، وتقديم مجموعات بيانات مهيكلة (CSV، JSON، Google Sheets)
متطلبات إضافية
- خبرة عملية مع نماذج اللغة الكبيرة وأطر الذكاء الاصطناعي لتعزيز الأتمتة وحل المشكلات
- انتباه قوي للتفاصيل والالتزام بدقة البيانات
- أخلاقيات عمل ذاتي التوجيه والقدرة على استكشاف المشكلات بشكل مستقل
- رابط إلى GitHub يعتبر إضافة
- إتقان الإنجليزية: مستوى فوق المتوسط (B2) أو أعلى (المطلوب)
بالنسبة لهذا المشروع، من المتوقع أن تتطلب المهام حوالي 10–20 ساعة أسبوعياً خلال المراحل النشطة، بناءً على متطلبات المشروع. هذا تقدير، ليس عبء عمل مضمون، وينطبق فقط أثناء نشاط المشروع.
في هذا المشروع، يمكن للمساهمين كسب حتى 37 دولاراً في الساعة مكافئاً، اعتماداً على المستوى وتيرة المساهمة. يتفاوت التعويض بين المشاريع بناءً على النطاق والتعقيد والخبرة المطلوبة. يرجى ملاحظة أن مشاريع أخرى على المنصة قد تقدم مستويات كسب مختلفة بناءً على متطلباتها.
الملف المرشح المطلوب
المؤهلات التعليمية
خبرة ذات صلة لا تقل عن 3 سنوات في هندسة البيانات، كشط الويب، الأتمتة، أو تطوير البرمجيات (مطلوب).
درجة البكالوريوس أو الماجستير في الهندسة، الرياضيات التطبيقية، علوم الكمبيوتر، أو مجالات تقنية ذات صلة تعتبر ميزة.
الخبرة الأكاديمية و/أو المهنية
ينبغي أن يمتلك المرشحون أساساً تقنياً قوياً وخبرة عملية في البرمجة النصية، الأتمتة، وتدفقات استخراج البيانات. نحن نبحث عن متخصصين قادرين على حل مشكلات غير تافهة، والعمل بثقة مع تقنيات الويب الحديثة وأدوات معالجة البيانات، وجمع وتنسيق والتحقق من البيانات من مصادر متنوعة بشكل منهجي. نهج منضبط وتفصيلي والقدرة على العمل بشكل مستقل أمران أساسيان.