Role OverviewWe're building a dataset to evaluate AI coding agents by creating challenging tasks and evaluation criteria within realistic simulated environments. You'll design tasks, write tests, and iterate on tasks and tests based on QA feedback.
What You Will Do
Create tasks and evaluation criteria, write tests, and iterate on tasks and tests based on QA feedback to evaluate AI coding agents.
Why It Might Be a Fit
You'll need to have 5+ years in software development, experience writing tests, and a strong understanding of AI and machine learning. You'll also need to have a B2+ level of English proficiency.
Requirements
- 5+ years in software development
- Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
- Experience writing tests (functional, integration)
- English proficiency - B2+
Benefits
- Up to $50/hr equivalent, depending on level and pace
- Tasks are estimated at ~20 hours each; you set your own schedule
]]>