AI Safety Specialist

Mercor
New York, NY
Job Description
Role Overview

The AI Safety Specialist will work independently and asynchronously to identify jailbreaks, prompt injections, and misuse cases in conversational AI models and agents. They will generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks. The Specialist will apply structure by following taxonomies, benchmarks, and playbooks to ensure consistent testing and document reproducibly by producing reports, datasets, and attack cases that customers can act on.

What You Will Do

The main day-to-day responsibilities will include red teaming conversational AI models and agents, generating human data, applying structure, and documenting results. The Specialist will work independently and asynchronously to meet deadlines while improving AI model performance.

Why It Might Be a Fit

The ideal candidate will have prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing, and be able to explain risks clearly to technical and non-technical stakeholders. Experience in Adversarial ML, Cybersecurity, or socio-technical risk analysis is preferred, as well as skills in psychology, acting, or writing for unconventional adversarial thinking.

Requirements

  • Fluent in English & Punjabi with native fluency
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Ability to explain risks clearly to technical and non-technical stakeholders
]]>