Role OverviewWe're building a dataset to evaluate AI coding agents by creating challenging tasks and evaluation criteria within realistic simulated environments. You'll design tasks, write tests, and iterate on tasks and tests based on QA feedback.
What You Will Do
Create tasks and evaluation criteria, write tests, and iterate on tasks and tests based on QA feedback. You'll guide and evaluate AI agent solutions, analyzing failures and refining until the evaluation is fair and robust.
Why It Might Be a Fit
You need to have 5+ years in software development, experience writing tests, and English proficiency (B2+). You'll work with a core stack of Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, and Redis.
Requirements
- 5+ years in software development
- Experience writing tests (functional, integration)
- English proficiency (B2+)
Benefits
- Up to $50/hr equivalent
- Flexible schedule
- Estimated 20 hours per task
]]>