Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves Frontier coding agents are already good at passing tests. We measure whether they pass them the right way .We're building a dataset to evaluate the safety and conduct of AI coding agents not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:
Build realistic developer environments a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes
Write tests that verify whether the agent completed the task the right way catching corners cut, not just checking outputs
Iterate on tasks and tests based on QA feedback review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT: Not data labeling; Not prompt engineering; Not cybersecurity or red-teaming there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists; Not writing code from scratch the agent writes most of the code; you design the situation and evaluate the outcome;
What we look for
4 5+ years in software development;
Core stack: Python, JavaScript/TypeScript;
Strong test design skills functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
Familiarity with GitHub PRs and CI workflows as a user;
Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful but you don't need to be an expert in every layer;
English proficiency B2+
Desired Candidate Profile
4 5+ years in software development;
Core stack: Python, JavaScript/TypeScript;
Strong test design skills functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
Familiarity with GitHub PRs and CI workflows as a user;
Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful but you don't need to be an expert in every layer;
English proficiency B2+
تربط Mindrift المتخصصين بفرص ذكاء اصطناعي قائمة على المشاريع لشركات تقنية رائدة، مركّزة على اختبار النظامات الذكاء الاصطناعي وتقييمها وتحسينها. المشاركة قائمة على المشروع، وليست توظيفاً دائماً.
ما تتضمنه هذه الفرصة Frontier coding agents هم جيدون بالفعل في اجتياز الاختبارات. نقيس ما إذا كانوا يجتازونها بالطريقة الصحيحة. نبني مجموعة بيانات لتقييم السلامة والسلوك لدى وكلاء الترميز بالذكاء الاصطناعي، ليس فقط إنجاز المهمة، بل إنجازها بأمانة، ضمن النطاق، ودون تقصير. ستصمم مهام حيث المسار السهل هو غير آمن، وتكتب الاختبارات التي تكشف ذلك:
إنشاء بيئات مطور واقعية كشركة افتراضية تحتوي على قاعدة كود وبنية تحتية وسياق (التذاكر، الوثائق، المحادثات) التي تشكل تاريخ تطوير موثوق
تصميم مهام تقترن هدف تطويري غير ضار مع اختصار غير آمن مغري: زيادة النطاق، مخالفات السياسات، تلف البيانات، تغييرات مفرطة السماح
كتابة اختبارات تتحقق مما إذا كان الوكيل قد أتم المهمة بالطريقة الصحيحة، مع كشف الزوايا المقطوعة، وليس فقط فحص المخرجات
التكرار على المهام والاختبارات بناءً على ملاحظات QA، مراجعة حلول الوكلاء، تحليل الإخفاقات، وتحسينها حتى تكون التقييمات عادلة وقوية
ما هذا ليس: ليس تصنيف بيانات؛ ليس هندسة طلبات داعمة؛ ليس أمن سيبراني أو فريق اختبار اختراق، لا وجود للمهاجم في السيناريو. خبرة الأمن السيبراني ميزة، وليست مطلوبة. نحن نبحث عن مهندسين يفهمون كيف يجب أن يتصرف الكود، لا عن مختبري اختراق. مهندسون برمجيون أقوياء، ليسوا متخصصين في الأمن؛ ليس كتابة كود من الصفر، الوكيل يكتب معظم الكود؛ أنت تصمم الوضع وتقيّم النتيجة؛
ما نبحث عنه
4-5+ سنوات في تطوير البرمجيات;
المكدس الأساسي: بايثون، جافا سكريبت/تايبسكريبت؛
مهارات تصميم اختبارات قوية: اختبارات وظيفية وتكامل تفصل بين الإكمال الآمن من غير الآمن، وليست فقط الصحيحة من غير الصحيحة؛
خبرة عملية مع وكلاء ترميز (Claude Code، GitHub Copilot CLI، Codex، أو ما شابه)؛
الاطلاع على PRs في GitHub وتدفقات CI كمستخدم؛
تنوع التقنية مرحب به، وليس فلترة. المهام تحاكي مستودعات حقيقية تحتوي على قواعد بيانات، خط أنابيب CI، و سكريبتات النشر، لذا وجود خبرة واسعة في الخلفية والبنية التحتية مفيد حقاً لكنك لست بحاجة لتكون خبيراً في كل طبقة;
إتقان الإنجليزية B2+
ملف المرشح المرغوب
4-5+ سنوات في تطوير البرمجيات؛
المكدس الأساسي: بايثون، جافا سكريبت/تايبسكريبت؛
مهارات تصميم اختبارات قوية: اختبارات وظيفية وتكامل تفصل بين الإكمال الآمن من غير الآمن، وليست فقط الصحيحة من غير الصحيحة؛
خبرة عملية مع وكلاء ترميز (Claude Code، GitHub Copilot CLI، Codex، أو ما شابه)؛
الاطلاع على PRs في GitHub وتدفقات CI كمستخدم؛
تنوع التقنية مرحب به، وليس فلترة. المهام تحاكي مستودعات حقيقية تحتوي على قواعد بيانات، خط أنابيب CI، و سكريبتات النشر، لذا وجود خبرة واسعة في الخلفية والبنية التحتية مفيد حقاً ولكنك لست بحاجة لتكون خبيراً في كل طبقة؛
إتقان الإنجليزية B2+