Senior Python Engineer - AI Coding Agent Evaluation (Freelance)

Mindrift
Dallas, TX
Remote
Job Description
Role Overview

We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria within realistic simulated environments.

What You Will Do

Create tasks and evaluation criteria, write tests, and iterate on tasks and tests based on QA feedback.

Why It Might Be a Fit

You'll need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.

Requirements

  • 8+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Benefits

  • Paid per accepted task
  • Rate depends on qualification tier and efficiency
  • Up to the equivalent of $200/hr
]]>