Role OverviewWe're building a dataset to evaluate AI coding agents by creating challenging tasks and evaluation criteria within realistic simulated environments. You'll design tasks, write tests, and iterate on feedback to ensure the evaluation is fair and robust.
What You Will Do
You'll create tasks from intermediate states of environments, write tests to verify agent solutions, and iterate on tasks and tests based on QA feedback.
Why It Might Be a Fit
We look for 5+ years of software development experience, proficiency in Python, JavaScript/TypeScript, Docker, Postgres, Kafka, and Redis, and experience writing tests (functional, integration).
Requirements
- 5+ years in software development
- Proficiency in Python, JavaScript/TypeScript, Docker, Postgres, Kafka, and Redis
- Experience writing tests (functional, integration)
- English proficiency - B2+
Benefits
- Up to $50/hr equivalent
- Flexible schedule
- 20 hours per task
]]>