AI Safety Expert - Red Team

Mercor
New York, NY
Job Description
Role Overview

As an AI Safety Expert - Red Team, you will be responsible for conversational AI models and agents to identify jailbreaks, prompt injections, and misuse cases. You will generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks. You will apply structure by following taxonomies, benchmarks, and playbooks to ensure consistent testing. You will document reproducibly to produce reports, datasets, and attack cases that customers can act on.

What You Will Do

Your main day-to-day responsibilities will include reviewing AI outputs on sensitive topics such as bias and misinformation, with optional participation in higher-sensitivity projects.

Why It Might Be a Fit

You will be a fit for this role if you have prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing, and can explain risks clearly to technical and non-technical stakeholders.

Requirements

  • Fluent Language Skills Required: English & Malay
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to explain risks clearly to technical and non-technical stakeholders

Benefits

  • $17–$25/hour
]]>