AI Evaluation Engineer (Python, QA or Security)

Mindrift
San Antonio, TX
Remote
Job Description
Role Overview

We're building a dataset to evaluate AI coding agents. You'll create challenging tasks and evaluation criteria within realistic simulated environments.

What You Will Do

Create tasks, design evaluation criteria, write tests, and iterate on tasks and tests based on QA feedback.

Why It Might Be a Fit

You need to have 5+ years in software development, experience writing tests, and be proficient in English (B2+).

Requirements

  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Benefits

  • Up to $50/hr equivalent
  • Flexible schedule
  • 20 hours per task, estimated
]]>