AI Safety Expert - Red Team

Mercor
San Francisco, CA
Job Description
Role Overview

AI Safety Expert - Red Team will be responsible for conversational AI models and agents to identify jailbreaks, prompt injections, and misuse cases. The role involves generating high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks. The expert will apply structure by following taxonomies, benchmarks, and playbooks to ensure consistent testing and document reproducibly to produce reports, datasets, and attack cases that customers can act on.

What You Will Do

Conversational AI models and agents, generate human data, apply structure, document reproducibly, review AI outputs on sensitive topics.

Why It Might Be a Fit

The ideal candidate will have prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing, ability to explain risks clearly to technical and non-technical stakeholders, and preferred experience in Adversarial ML, background in Cybersecurity, and expertise in socio-technical risk and creative probing skills.

Requirements

  • Fluent Language Skills Required: English & Malay
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to explain risks clearly to technical and non-technical stakeholders
]]>