Adversarial ML Specialist

Mercor
San Francisco, CA
Category Security
Job Description
Role Overview

As an Adversarial ML Specialist, you will work independently and asynchronously to red team conversational AI models and agents, conduct jailbreaks, and identify vulnerabilities. You will also generate high-quality human data and document reproducibly by producing reports, datasets, and attack cases for customer action.

What You Will Do

Your main day-to-day responsibilities will include red teaming, generating human data, annotating failures, classifying vulnerabilities, and flagging systemic risks. You will also apply structure using taxonomies, benchmarks, and playbooks to maintain consistent testing.

Why It Might Be a Fit

To be a fit for this role, you must have prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing, and strong communication skills to explain risks to technical and non-technical stakeholders.

Requirements

  • Fluent in English and Swedish
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing
  • Strong communication skills to explain risks to technical and non-technical stakeholders
]]>