Role OverviewAs an AI Red Team Specialist, you will probe models to explore how frontier AI models handle coding, ML, and analysis tasks, identifying spots where models quietly fail. You will design challenges by turning identified weaknesses into well-crafted tasks that are challenging for models but fair to grade. You will document findings clearly with reproducible evidence and steps, and strengthen tasks by collaborating with task authors to close loopholes, shortcuts, and grading gaps.
What You Will Do
Your main day-to-day responsibilities will include probing models, designing challenges, documenting findings, and collaborating with task authors to improve the benchmark.
Why It Might Be a Fit
This role might be a fit for you if you have a strong background in research, research-engineering, security, or AI-evaluation roles, and you are able to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems.
Requirements
- MSc or PhD in a STEM field or equivalent practical experience
- 1+ years of experience in research, research-engineering, security, or AI-evaluation roles
- Ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems
- Proficiency in Python and Git for scripting probes and analyses
- Familiarity with LLM capabilities, limitations, and evaluation techniques
- Ability to engage reliably for approximately 35 hours/week
Benefits
]]>