Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
About the Role
You ll design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome. Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.
Responsibilities :
- Invent a realistic developer scenario a real bug, a broken ETL, a missing feature not a toy problem.
- Build a reproducible Docker environment with pinned dependencies.
- Write a pytest that verifies outcomes, not specific commands deterministic, non-flaky, and does not leak the fix.
- Write an instruction.md that reads like a Jira ticket a developer would receive.
- Write a reference solve.sh proving the task is solvable.
- Calibrate difficulty so current state-of-the-art agents solve the task 20 60% of the time.
- Iterate based on feedback from expert QA reviewers.
- Later: review other authors tasks as a QA reviewer.
Not in scope
Data labeling, prompt engineering. Production code to ship you design problems and verification for AI agents. Leetcode puzzles scenarios must look like real developer work. Not every candidate task ships quality over quantity.
Requirements
- 3+ years of production software development in one backend stack Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth.
- Python + pytest fluency required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py.
- Docker authoring reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.
- Linux & Bash comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.
- AI coding agent experience Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it.
- English B2+ written.
Not a fit
Data Science, ML, or Computer Vision engineers without backend-engineering output. Manual QA testers without automation or test authoring. Frontend-only, low-code / no-code, IT Support, or Business Analysts. Engineers who have never written pytest from scratch. Junior, intern, or assistant as the most recent role.
Preferred qualifications
Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals. Modern Python tooling (uv, poetry, pyproject.toml). Coverage tooling (pytest-cov, coverage.py, gcov, llvm-cov, kcov). Fuzzing or property-based testing (Hypothesis). Prior contribution to agent-evaluation benchmarks or related frameworks.
Process
Apply Pass qualification (90-minute sample-task screen + short behavioral interview) Join a project Complete tasks Get paid.
Time commitment
Onboarding: ~10 hours per first task. Steady state: ~5 hours per task, 2 4 parallel tasks per author. Realistic weekly load: 8 20 hours. Higher volume available for top performers. You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.
Compensation: Paid contributions, rates up to $35/hour *. Task-based compensation equivalent to hourly rate, depending on performance and volume. Some projects include incentive payments. *Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.
Desired Candidate Profile
- 3+ years of production software development in one backend stack Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth.
- Python + pytest fluency required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py.
- Docker authoring reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.
- Linux & Bash comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.
- AI coding agent experience Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it.
- English B2+ written.
Not a fit
Data Science, ML, or Computer Vision engineers without backend-engineering output. Manual QA testers without automation or test authoring. Frontend-only, low-code / no-code, IT Support, or Business Analysts. Engineers who have never written pytest from scratch. Junior, intern, or assistant as the most recent role.
Preferred qualifications
Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals. Modern Python tooling (uv, poetry, pyproject.toml). Coverage tooling (pytest-cov, coverage.py, gcov, llvm-cov, kcov). Fuzzing or property-based testing (Hypothesis). Prior contribution to agent-evaluation benchmarks or related frameworks.
تربط Mindrift المتخصصين بفرص الذكاء الاصطناعي القائمة على المشاريع لشركات التكنولوجيا الرائدة، مع التركيز على اختبار وتقييم وتحسين أنظمة الذكاء الاصطناعي. المشاركة قائمة على المشروع وليست توظيفاً دائماً.
عن الدور
ستصمم مهام ترميز تتحدى وكلاء ترميز AI في الطرف الأمامي. كل مهمة هي بيئة Docker مكتملة بذاتها مع قطعة برمجية مكسورة؛ يحاول وكيل AI إصلاحها؛ الاختبارات الآلية تتحقق من النتيجة. المنتج النهائي المتوقع هو حزمة المهمة الكاملة: الشفرة المكسورة، الاختبارات، التعليمات، وحل مرجعي يثبت أن المهمة قابلة للحل.
المسؤوليات :
- ابتكار سيناريو مطور واقعي—خلاله علة حقيقية، ETL معطّل، ميزة مفقودة وليس مجرد مشكلة لعبة.
- بناء بيئة Docker قابلة لإعادة الإنتاج مع الاعتمادات المؤكدة.
- كتابة pytest يتحقق من النتائج، وليس أوامر محددة؛ حتمي، غير flaky، ولا يكشف عن الإصلاح.
- كتابة instruction.md يقرأ كتذكرة Jira يتلقاها مطور.
- كتابة حل مرجعي solve.sh يثبت أن المهمة قابلة للحل.
- معايرة الصعوبة بحيث يحلها وكلاء state-of-the-art الحاليون بنجاح 20-60% من الوقت.
- الت iterate بناءً على تعليقات مراجعي QA الخبراء.
- في وقت لاحق: مراجعة مهام المؤلفين الآخرين كمراجع QA.
غير ضمن النطاق
تصنيف البيانات، وتنسيق الأوامر. كود الإنتاج لإرسالك صُمم مشاكل التصميم والتحقق لوكلاء AI. سيناريوهات ألغاز Leetcode يجب أن تبدو كعمل مطور حقيقي. ليست كل مهمة مرشحة للجودة العالية مع الكمية العالية.
المتطلبات
- 3+ سنوات من تطوير البرمجيات الإنتاجية في سلسلة خلفية واحدة Python، Go، Node.js، Java، أو Rust. العمق في سلسلة واحدة يتحسن على حساب الاتساع.
- إتقان Python + pytest مطلوب بغض النظر عن السلسلة الأساسية. الحزمة التجريبية قائمة على pytest حتى وإن كانت التطبيق المكسور بلغة أخرى. Fixtures، parametrize، monkeypatch، timeouts، conftest.py.
- كتابة Docker ملفات قابلة لإعادة الإنتاج، اعتمادات مثبتة، بناء متعدد المراحل عند الحاجة، مستخدم غير جذر.
- راحة Linux & Bash في التصحيح داخل الحاويات (strace، lsof، journalctl)؛ shell بما يتجاوز set -euo pipefail.
- خبرة وكيل ترميز AI مثل Claude Code، Cursor، Roo Code، أو ما شابهها، في عمل غير تافه. يمكنك ذكر وقت محدد كان فيه الذكاء الاصطناعي مخطئًا بثقة وكيف اكتشفت ذلك.
- كتابة باللغة الإنجليزية B2+ مكتوبة.
غير مناسب
مهندسو البيانات، ML، أو الرؤية الحاسوبية بدون نتائج للهندسة الخلفية. مختبرو QA يدويون بدون أتمتة أو تأليف اختبارات. مطورون أماميون فقط، أو بدون كود، أو دعم تكنولوجي، أو محللو أعمال. مهندسون لم يكتبوا pytest من الصفر. مبتدئ، متدرّب، أو مساعد كأقرب دور.
المؤهلات المفضلة
عمق في الأمن، إدارة النظام (nginx / systemd / cron)، الحوسبة العلمية (NumPy / PyTorch / SciPy)، DevOps، أو أسرار Git. أدوات Python حديثة (uv، poetry، pyproject.toml). أدوات التغطية (pytest-cov، coverage.py، gcov، llvm-cov، kcov). fuzzing أو اختبار قائم على الخواص (Hypothesis). مساهمة سابقة في معايير تقييم الوكلاء أو أطر ذات صلة.
العملية
التقديم تمرير تأهيل Pass (90 دقيقة عينة مهمة + مقابلة سلوكية قصيرة) انضم لمشروع أكمل المهام ادفع.
الالتزام الزمني
التوجيه: ~10 ساعات للمهمة الأولى. الحالة المستقرة: ~5 ساعات لكل مهمة، 2-4 مهام بالتوازي لكل مؤلف. تحميل أسبوعي واقعي: 8-20 ساعة. حجم أعلى متاح لأصحاب الأداء العالي. اختر متى وكيف تساهم؛ يجب تقديم المهام قبل الموعد النهائي وتفي بمعايير القبول.
التعويض: مساهمات مدفوعة الأجر، حتى 35$/ساعة *. تعويض قائم على المهمة يعادل الأجر بالساعة، حسب الأداء والحجم. بعض المشاريع تشمل دفعات حافزة. *تتفاوت الأسعار بناءً على الخبرة وتقييم المهارات والموقع واحتياجات المشروع وعوامل أخرى. قد تُقدم أسعار أعلى للخبراء المتخصصين جدًا. قد تُطبق أسعار أدنى أثناء الإندماج أو مراحل المشروع غير الأساسية. يتم مشاركة تفاصيل المدفوعات لكل مشروع.
المرشح المنشود
- 3+ سنوات من تطوير البرمجيات الإنتاجية في واحدة من سلاسل الخلفية Python، Go، Node.js، Java، أو Rust. العمق في سلسلة واحدة يفوق الاتساع.
- إتقان Python + pytest مطلوب بغض النظر عن السلسلة الأساسية. الحزمة التجريبية قائمة على pytest حتى وإن كانت التطبيق المكسور بلغة أخرى. Fixtures، parametrize، monkeypatch، timeouts، conftest.py.
- كتابة Docker ملفات قابلة لإعادة الإنتاج، اعتمادات مثبتة، بناء متعدد المراحل عند الحاجة، مستخدم غير جذر.
- راحة Linux & Bash في التصحيح داخل الحاويات (strace، lsof، journalctl)؛ shell بما يتجاوز set -euo pipefail.
- خبرة وكيل ترميز AI مثل Claude Code، Cursor، Roo Code، أو ما شابهها، في عمل غير تافه. يمكنك ذكر وقت محدد كان فيه الذكاء الاصطناعي مخطئًا بثقة وكيف اكتشفت ذلك.
- كتابة باللغة الإنجليزية B2+ مكتوبة.
غير مناسب
مهندسو البيانات، ML، أو الرؤية بالحاسوب بدون نتائج للهندسة الخلفية. مختبرو QA يدويون بدون أتمتة أو تأليف اختبارات. مطورون أماميون فقط، أو بدون كود، أو دعم تكنولوجي، أو محللو أعمال. مهندسون لم يكتبوا pytest من الصفر. مبتدئ، متدرّب، أو مساعد كأقرب دور.
المؤهلات المفضلة
عمق في الأمن، إدارة النظام (nginx / systemd / cron)، الحوسبة العلمية (NumPy / PyTorch / SciPy)، DevOps، أو أسرار Git. أدوات Python حديثة (uv، poetry، pyproject.toml). أدوات التغطية (pytest-cov، coverage.py، gcov، llvm-cov، kcov). fuzzing أو اختبار قائم على الخواص (Hypothesis). مساهمة سابقة في معايير تقييم الوكلاء أو أطر ذات صلة.