Role OverviewWe're building a dataset to evaluate AI coding agents by creating challenging tasks and evaluation criteria within realistic simulated environments.
What You Will Do
You'll create tasks, design evaluation criteria, write tests, and iterate on tasks and tests based on QA feedback.
Why It Might Be a Fit
You need 5+ years in software development, experience writing tests, and English proficiency (B2+).
Requirements
- 5+ years in software development
- Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
- Experience writing tests (functional, integration)
- English proficiency - B2+
Benefits
- Up to $50/hr equivalent
- Flexible schedule
- 20 hours per task, estimated
]]>