AI Safety Expert - Red Team

Mercor
San Francisco, CA
Job Description
Role Overview

As an AI Safety Expert, you will work on red teaming conversational AI models and agents, focusing on jailbreaks, prompt injections, misuse cases, and bias exploitation. You will generate high-quality human data, annotate failures, classify vulnerabilities, and flag systemic risks. You will apply structure by following taxonomies, benchmarks, and playbooks to maintain testing consistency.

What You Will Do

Your day-to-day responsibilities will include generating high-quality human data, annotating failures, classifying vulnerabilities, and flagging systemic risks. You will also apply structure by following taxonomies, benchmarks, and playbooks to maintain testing consistency.

Why It Might Be a Fit

To be a fit for this role, you should have prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing. You should also have strong communication skills to explain risks clearly to technical and non-technical stakeholders.

Requirements

  • Fluent in English and Malay
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Strong communication skills to explain risks clearly to technical and non-technical stakeholders

Benefits

  • $17–$25/hour
]]>