AI Safety Expert - Red Team

Mercor
New York, NY
Job Description
Role Overview

The AI Safety Expert - Red Team will work independently and asynchronously to identify jailbreaks, prompt injections, misuse cases, and bias exploitation in conversational AI models and agents. They will generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks. The expert will apply structure by following taxonomies, benchmarks, and playbooks to ensure consistent testing and document reproducibly by producing reports, datasets, and attack cases that customers can act on.

What You Will Do

The main day-to-day responsibilities of the AI Safety Expert - Red Team will include conversational AI model testing, human data generation, and documentation of findings. They will work to improve AI model performance and identify potential risks and vulnerabilities.

Why It Might Be a Fit

The ideal candidate will have prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing, as well as strong communication skills to explain risks to technical and non-technical stakeholders. Experience with Adversarial ML, Cybersecurity, and Socio-technical risk is preferred, as well as skills in creative probing such as psychology, acting, or unconventional adversarial thinking.

Requirements

  • Fluent in English and Swedish
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Strong communication skills to explain risks to technical and non-technical stakeholders
]]>