AI Evaluation Engineer (Python, QA or Security)

Mindrift
San Antonio, TX
Remote
Job Description
Role Overview

We're building a dataset to evaluate AI coding agents by creating challenging tasks and evaluation criteria within realistic simulated environments. You'll design tasks, write tests, and iterate on tasks and tests based on QA feedback.

What You Will Do

Create tasks and evaluation criteria, write tests, and iterate on tasks and tests based on QA feedback to ensure the evaluation is fair and robust.

Why It Might Be a Fit

You'll need 5+ years of software development experience, a core stack of Python, JavaScript, Docker, Postgres, Kafka, and Redis, and experience writing tests (functional, integration).

Requirements

  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Benefits

  • Up to $50/hr equivalent
  • Flexible schedule
  • Estimated 20 hours per task
]]>