Role OverviewAs an AI Workflow Evaluator, you will evaluate AI systems on complex personal workflows, design realistic prompts, and write clear explanations of AI successes and failures to improve AI model performance.
What You Will Do
Your day-to-day responsibilities will include evaluating AI outputs for practicality, personalization, and reasoning, identifying where models miss context, overreach, or fail to use tools correctly, and writing against rubrics.
Why It Might Be a Fit
This role requires strong written judgment, attention to detail, and experience using LLM plugins/connectors, as well as a willingness to sign a data-share consent form and evaluate AI outputs against rubrics.
Requirements
- US-based only
- Strong MCP experience and plugin/connector usage
- Experience using LLM plugins/connectors like Google Drive, Expedia, and Notion
- Heavy personal usage of LLM products
- Active LLM account with 6+ months of history
- Willingness to sign a data-share consent form via DocuSign
- Experience using AI for multi-step planning, research, and decision-making
- Strong written judgment and attention to detail
- Ability to evaluate AI outputs and write against rubrics
Benefits
- Hourly compensation: $50–$190/hour
]]>