Software Engineering Evaluation Specialist

Mindrift
San Antonio, TX
Remote
Job Description
Role Overview

You'll design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome. Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.

What You Will Do

Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature — not a toy problem. Build a reproducible Docker environment with pinned dependencies. Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix.

Why It Might Be a Fit

You have 3+ years of production software development in one backend stack — Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth. You're fluent in Python + pytest and have experience with Docker authoring, Linux & Bash, and AI coding agents.

Requirements

  • 3+ years of production software development in one backend stack — Python, Go, Node.js, Java, or Rust
  • Python + pytest fluency
  • Docker authoring
  • Linux & Bash
  • AI coding agent experience
  • English — B2+ written

Benefits

  • Paid contributions, rates up to $35/hour*
  • Task-based compensation equivalent to hourly rate, depending on performance and volume
  • Some projects include incentive payments
]]>