Mindrift

AI Evaluation Engineer (Python, QA, Security)

Mindrift

Qatar · Part Time

Be the first to apply

Experience
5+ yrs
Salary
USD 50 – USD 50 / hour
Openings
1
Posted
2 hours ago
Work mode
In office
Eligibility
Open to candidates with a minimum of five years in software development and English proficiency of B2 or higher.
Resume
Required to apply

Job description

Overview

Mindrift connects experts with project-oriented AI jobs conducted for leading technology companies, focusing on evaluating, testing, and improving AI systems. These are project-based contracts rather than permanent positions. The current project involves creating a dataset to assess AI coding agents' capabilities in handling realistic software development challenges.

Key Responsibilities

  • Develop realistic simulated developer environments reflecting a virtual company: including a codebase, infrastructure, and contextual information such as tickets, documents, and conversations depicting authentic development history.
  • Create tasks starting from intermediate states in these environments by designing prompts, defining clear success criteria, and ensuring tasks are solvable by AI agents.
  • Develop verification tests that validate AI agent solutions, accommodating all valid approaches while rejecting incorrect ones; aiming for balanced test strictness.
  • Continuously refine tasks and tests through quality assurance feedback by analyzing AI failures and improving the evaluation framework for fairness and robustness.

What This Role Does Not Involve

  • This position is not about data labeling.
  • It does not require prompt engineering.
  • Coding from scratch is not expected as AI agents primarily write the code, while you guide and assess their outputs.

Candidate Profile

  • Minimum 5 years of professional software development experience.
  • Expertise in Python (particularly FastAPI), JavaScript/TypeScript (React), Docker, PostgreSQL, Kafka, and Redis.
  • Proven experience in writing various test types including functional and integration tests.
  • English proficiency at B2 level or higher is required.

Challenges

Creating genuinely challenging tasks for frontier AI coding models requires deep insight into AI failure modes and the capacity to identify scenarios that distinguish between high- and low-quality AI-generated code. Developing tests that correctly accept all valid solutions but reject poor ones is especially complex.

Process

Applicants go through qualification steps, join projects, complete assignments, and receive payments based on performance.

Compensation and Scheduling

Payments can reach up to $50 per hour, contingent on your expertise and working pace. Typical tasks require about 20 hours. You control your own schedule within the project scope.

Additional Information

Please submit your CV in English and specify your English proficiency level.

Work styles they’re looking for

Analytical Thinking Attention to Detail

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help