Job Description
Location n
Remote
Fluent Language Skills Required n
English & Kannada. Native fluency in English and Kannada is required for this position.
Why This Role Exists n
At Mercor, we believe the safest AI is the one that's already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers.
n
This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated.
What You'll Do n
n
n
Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation
n
n
Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks
n
n
Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent
n
n
Document reproducibly: produce reports, datasets, and attack cases customers can act on
n
Who You Are n
n
n
You bring strong judgment about language and content: you can tell whether an AI response is accurate, complete, and appropriate, and explain why
n
n
You're rigorous: you notice subtle errors, inconsistencies, and gaps that others skim past
n
n
You're structured: you work to guidelines and quality standards consistently, not ad hoc
n
n
You're communicative: you explain your reasoning clearly to technical and non-technical audiences
n
n
You're adaptable: you thrive moving across projects, task types, and customers
n
Nice-to-Have Specialties n
n
n
Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction
n
n
Cybersecurity: penetration testing, exploit development, reverse engineering
n
n
Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing
n
n
Creative probing: psychology, acting, writing for unconventional adversarial thinking
n
What Success Looks Like n
n
n
You uncover vulnerabilities automated tests miss
n
n
You deliver reproducible artifacts that strengthen customer AI systems
n
n
Evaluation coverage expands: more scenarios tested, fewer surprises in production
n
n
Mercor customers trust the safety of their AI because you've already probed it like an adversary
n
Why Join Mercor n
n
n
Build experience in human data-driven AI red teaming at the frontier of safety
n
n
Play a direct role in making AI systems more robust, secure, and trustworthy
n
n
#J-*****-Ljbffr
📌 Ai Safety Specialist - Fully Remote (Perth)
🏢 Mercor
📍 Perth