Role OverviewAI Evaluation Engineer will create challenging tasks and evaluation criteria for AI systems, building realistic developer environments and designing tasks from intermediate states of these environments. They will write tests that verify agent solutions and iterate on tasks and tests based on QA feedback.
What You Will Do
Create tasks, write tests, and evaluate AI agent solutions in a project-based setting. Tasks involve building realistic developer environments, designing tasks from intermediate states, and writing tests to verify agent solutions.
Why It Might Be a Fit
This role requires 5+ years of software development experience, with a core stack of Python, JavaScript/TypeScript, Docker, Postgres, Kafka, and Redis. Experience in writing tests and English proficiency (B2+) are also required.
Requirements
- 5+ years in software development
- Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
- Experience writing tests (functional, integration)
- English proficiency - B2+
Benefits
- Up to $50/hr equivalent
- Flexible schedule
- Project-based work
]]>