Role OverviewAI Evaluation Engineer will create challenging tasks and evaluation criteria within realistic simulated environments to evaluate AI coding agents. The role involves building developer environments, designing tasks, writing tests, and iterating based on QA feedback.
What You Will Do
Create tasks, write tests, and iterate on tasks and tests based on QA feedback to ensure the evaluation is fair and robust. Tasks involve creating a believable development history, designing tasks from intermediate states, and writing tests that verify agent solutions.
Why It Might Be a Fit
The ideal candidate has 5+ years in software development, experience writing tests, and proficiency in Python, JavaScript/TypeScript, Docker, Postgres, Kafka, and Redis. The role requires deep understanding of where models fail and what scenarios reveal the difference between a good and a bad solution.
Requirements
- 5+ years in software development
- Experience writing tests (functional, integration)
- Proficiency in Python, JavaScript/TypeScript, Docker, Postgres, Kafka, and Redis
- English proficiency - B2+
Benefits
- Up to $50/hr equivalent
- Flexible schedule
- Tasks are estimated at ~20 hours each
]]>