04 Sep
|
Mercor
|
New South Wales
04 Sep
Mercor
New South Wales
Job Description
Location: Remote
n
Fluent Language Skills Required: English & Odia. Native fluency in English and Odia is required for this position.
n
Why This Role Exists
n
At Mercor, we believe the safest AI is the one that is already been attacked - by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers.
n
This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated.
n
What You'll Do
n
n
- Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation
n
- Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks
n
- Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent
n
- Document reproducibly: produce reports, datasets, and attack cases customers can act on
n
n
Who You Are
n
n
- You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing)
n
- You're curious and adversarial: you instinctively push systems to breaking points
n
- You're structured: you use frameworks or benchmarks, not just random hacks
n
- You're communicative: you explain risks clearly to technical and non-technical stakeholders
n
- You're adaptable: thrive on moving across projects and customers
n
n
Nice-to-Have Specialties
n
n
- Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction
n
- Cybersecurity: penetration testing, exploit development, reverse engineering
n
- Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing
n
- Creative probing: psychology, acting, writing for unconventional adversarial thinking
n
n
What Success Looks Like
n
n
- You uncover vulnerabilities automated tests miss
n
- You deliver reproducible artifacts that strengthen customer AI systems
n
- Evaluation coverage expands: more scenarios tested, fewer surprises in production
n
- Mercor customers trust the safety of their AI because you've already probed it like an adversary
n
n
Why Join Mercor
n
n
- Build experience in human data-driven AI red teaming at the frontier of safety
n
- Play a direct role in making AI systems more robust, protected, and trustworthy
n
📌 AI Safety Expert - Red Team (New South Wales)
🏢 Mercor
📍 New South Wales