Software Engineering Evaluation Specialist

Mindrift
New York, NY
Remote
Job Description
Role Overview

Design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome.

What You Will Do

Invent a realistic developer scenario, build a reproducible Docker environment, write a pytest, and calibrate difficulty so current state-of-the-art agents solve the task 20–60% of the time.

Why It Might Be a Fit

You'll have the opportunity to work on project-based AI opportunities, design challenging tasks, and contribute to agent-evaluation benchmarks. You'll also have the chance to review other authors' tasks as a QA reviewer.

Requirements

  • 3+ years of production software development in one backend stack — Python, Go, Node.js, Java, or Rust.
  • Python + pytest fluency — required regardless of primary stack.
  • Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.
  • Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.
  • AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work.

Benefits

  • Paid contributions, rates up to $35/hour*.
  • Task-based compensation equivalent to hourly rate, depending on performance and volume.
  • Some projects include incentive payments.
]]>