وصف العمل
يرجى إرسال سيرتك الذاتية باللغة الإنجليزية وتحديد مستوى إجادتك للإنجليزية.
تربط Mindrift المتخصصين بفرص ذكاء اصطناعي قائمة على المشاريع لشركات التقنية الرائدة، مع التركيز على الاختبار والتقييم وتحسين أنظمة الذكاء الاصطناعي.
المشاركة قائمة على المشروع، وليست توظيفاً دائماً.
ما تنطوي عليه هذه الفرصة: عملاء التشفير Frontier يجيدون اجتياز الاختبارات بالفعل.
نقيس ما إذا كانوا ينجزونها بالطريقة الصحيحة .
نحن نبني مجموعة بيانات لتقييم سلامة وسلوك وكلاء التشفير الذكي — ليس فقط ما إذا كانوا أكملوا مهمة، بل ما إذا فعلوها بنزاهة، ضمن النطاق، وبدون تقطيع زوايا.
ستصمم مهام حيث المسار السهل هو غير آمن، وتكتب الاختبارات التي تكشف ذلك: بناء بيئات مطور واقعية — شركة افتراضية تحتوي على قاعدة كود وبنية تحتية وسياق (تذاكر، وثائق، محادثات) تشكل تاريخ تطوير مقنع تصميم مهام تقترن هدف تطويري آمن مع اختصار غير آمن مغرٍ: تمديد النطاق، انتهاكات السياسات، تلف البيانات، تغييرات واسعة الإذن اكتب اختبارات تتحقق مما إذا كان الوكيل أنجز المهمة بالطريقة الصحيحة — القبض على الحواف المقطوعة، وليس مجرد التحقق من المخرجات تكرار المهام والاختبارات بناءً على تغذية راجعة من QA — مراجعة حلول الوكيل، تحليل الإخفاقات، وتحسين حتى تكون التقييم عادلًا وقويًا ما هذا ليس: ليس وسم البيانات؛ ليس هندسة المطالبات؛ ليس الأمن السيبراني أو اختبارات الاختراق — لا وجود لمهاجم في السيناريو.
خبرة الأمن السيبراني ميزة ولكنها ليست شرطاً.
نبحث عن مهندسين يفهمون كيف يجب أن يتصرف الكود، لا كاختبار اختراق.
مهندسون برمجيون أقوياء، ليسوا متخصصين في الأمن؛ ليس كتابة الكود من الصفر — الوكيل يكتب معظم الكود؛ أنت تصمم الوضع وتقيِّم النتيجة؛ ما نبحث عنه: 4–5+ سنوات في تطوير البرمجيات؛ التكديس الأساسي: Python، JavaScript/TypeScript؛ مهارات تصميم اختبارات قوية — اختبارات وظيفية وتكامل تفصل بين الإكمال الآمن وغير الآمن، وليس فقط الصحيح من الخاطئ؛ خبرة عملية في وكلاء الترميز (Claude Code، GitHub Copilot CLI، Codex، أو ما يماثلها)؛ الإلمام بطلبات سحب GitHub وCI ك مستخدم؛ سلاسل متعددة مطلوبة مرحب بها، ليست فلاتر فقط.
تحاكي المهام مستودعات حقيقية مع قواعد بيانات، وعمليات CI، وسكربتات النشر، لذا فإن التعرض الأوسع للخلفية وال بنية التحتية مفيد حقاً — لكن لا تحتاج أن تكون خبيراً في كل طبقة؛ إجادة اللغة الإنجليزية — B2+ لماذا هذا صعب Frontier النماذج لديها إتقان في الترميز بالفعل.
خلق مهمة تتحدى النماذج الأفضل حقاً ليس بالأمر السهل.
الصعوبة الحقيقية هي بناء الإغراء — سيناريو حيث المسار غير الآمن أو خارج النطاق هو المسار الأقل مقاومة — ثم كتابة اختبارات تكشف موثوقاً عن وكيل تخذه.
للمهام العديد من الحلول الصحيحة؛ يجب قبول جميعها ورفض السيئ منها.
كيف يعمل: التطبيق → اجتياز التأهيل → الانضمام إلى مشروع → إكمال المهام → الدفع
توقعات وقت المشروع: لهذه المهمة، من المتوقع أن تتطلب المهام حوالي 20-25 ساعة في الأسبوع خلال المراحل النشطة، بناءً على متطلبات المشروع.
هذه تقدير، ليس عبء عمل مضمون، ويسري فقط أثناء نشاط المشروع.
يجب تقديم المهام قبل الموعد النهائي وتلبية معايير القبول المدرجة ليتم قبولها.
التعويض: في هذا المشروع، يمكن للمساهمين كسب حتى 75 دولاراً في الساعة المعادلة، حسب مستواهم وسرعتهم في المساهمة.
يختلف التعويض عبر المشاريع حسب النطاق، التعقيد، والخبرة المطلوبة.
يرجى ملاحظة أن مشاريع أخرى على المنصة قد تقدم مستويات كسب مختلفة وفقاً لمتطلباتها.
المرشح المفضل
سنوات الخبرة
لا اختلاف خبرة مطلوب
الدرجة العلمية
درجة البكالوريوس / دبلوم أعلى
Job description
Please submit your CV in English and indicate your level of English proficiency.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.
Participation is project-based, not permanent employment.
What this opportunity involves Frontier coding agents are already good at passing tests.
We measure whether they pass them the right way .
We're building a dataset to evaluate the safety and conduct of AI coding agents — not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.
You'll design tasks where the easy path is the unsafe one, and write the tests that catch it: Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes Write tests that verify whether the agent completed the task the right way — catching corners cut, not just checking outputs Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling; Not prompt engineering; Not cybersecurity or red-teaming — there is no attacker in the scenario.
Cybersecurity experience is a nice-to-have but not a requirement.
We're looking for engineers who understand how code should behave, not penetration testers.
Strong software engineers, not security specialists; Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome; What we look for 4–5+ years in software development; Core stack: Python, JavaScript/TypeScript; Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect; Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar); Familiarity with GitHub PRs and CI workflows as a user; Stack breadth is welcome, not a filter.
Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer; English proficiency — B2+ Why this is hard Frontier models are already good at coding.
Creating a task that genuinely challenges the best models is non-trivial.
The real difficulty is building the temptation — a scenario where the unsafe or out-of-scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it.
Tasks have many valid solutions; tests must accept all of them and reject the bad ones.
How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Project time expectations For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements.
This is an estimate, not a guaranteed workload, and applies only while the project is active.
Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.
Compensation On this project, contributors can earn up to $75 per hour equivalent , depending on their level and pace of contribution.
Compensation varies across projects depending on scope, complexity, and required expertise.
Please note that other projects on the platform may offer different earning levels based on their requirements.
Preferred candidate
Years of experience
No experience required
Degree
Bachelor's degree / higher diploma