AI Safety Expert - Red Team

Mercor
San Francisco, CA
Remote
Job Description
Role Overview

We are seeking an AI Safety Expert - Red Team to join our team in San Francisco. The successful candidate will be responsible for red teaming conversational AI models and agents, generating high-quality human data, and applying structured approaches to ensure consistent testing. The role requires strong communication skills to convey risks to technical and non-technical stakeholders.

What You Will Do

The AI Safety Expert - Red Team will be responsible for identifying jailbreaks, prompt injections, misuse cases, and bias exploitation, generating human data by annotating failures, classifying vulnerabilities, and flagging systemic risks, and documenting findings reproducibly to produce reports, datasets, and attack cases.

Why It Might Be a Fit

The ideal candidate will have prior experience in red teaming, AI adversarial work, cybersecurity, or socio-technical probing, and be able to communicate risks clearly to both technical and non-technical stakeholders.

Requirements

  • Fluent in English and Finnish
  • Prior experience in red teaming, AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to communicate risks clearly to both technical and non-technical stakeholders
]]>