Role OverviewAs a DevOps Engineer - AI Model Evaluator, you will use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks, review model-generated implementations, and apply professional engineering judgment to realistic infrastructure engineering scenarios.
What You Will Do
Your day-to-day responsibilities will include identifying bugs, edge cases, reliability issues, and failure modes, comparing outputs from multiple frontier models, and working independently and asynchronously to meet sprint-based project deadlines.
Why It Might Be a Fit
To be a fit for this role, you must have 2+ years of professional DevOps, SRE, or Cloud Engineering experience, and regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
Requirements
- 2+ years of professional DevOps, SRE, or Cloud Engineering experience
- Experience with AWS, Azure, GCP, Kubernetes, Terraform, CI/CD pipelines, or observability tooling
- Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools
- Ability to evaluate model-generated infrastructure and reliability engineering solutions
Benefits
- Compensation: $85/hour
- Compensation is tied to accepted work: $400 per accepted task
]]>