AI Evaluation Engineer (Python, QA or Security)

Mindrift
Austin, TX
Remote
Job Description
Role Overview

We're building a dataset to evaluate AI coding agents by creating challenging tasks and evaluation criteria within realistic simulated environments. You'll design tasks, write tests, and iterate on tasks and tests based on QA feedback.

What You Will Do

Create tasks and tests for AI agents, iterate on feedback, and refine the evaluation process to ensure it's fair and robust.

Why It Might Be a Fit

You'll need 5+ years of software development experience, proficiency in Python, JavaScript/TypeScript, Docker, Postgres, Kafka, and Redis, as well as experience writing tests and a strong understanding of where AI models fail.

Requirements

  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Benefits

  • Up to $50/hr equivalent, depending on level and pace
  • Tasks are estimated at ~20 hours each; you set your own schedule
]]>